Military aircraft were already in the air this spring when officials discovered that a planned U.S. strike on a Chinese vessel was based on intelligence an AI chatbot had hallucinated, CNN reported, citing U.S. officials. Armed personnel had been preparing to board the ship in the Middle East before the operation was aborted at the last minute.

A Special Operations Command analyst had used the chatbot to combine open source data with classified signals intelligence, producing a report that claimed the vessel was carrying components for a nuclear weapons program, according to CNN. The report circulated during the U.S. war with Iran and was, in the words of CNN's sources, entirely false. The operation was aborted at the last minute.

Jake Steckler, a research scholar at GovAI and a former U.S. Army officer, told TechCrunch that service members need to understand "the uncertainty inherent to LLMs," especially for "decisions that could lead to use of force, like targeting, intelligence analysis, or operational planning," because "there are life and death consequences for those decisions," he said.

The failure here wasn't a model refusing a request or producing something offensive. It was a fabricated fact that read as credible enough to move up a chain of command until someone caught it minutes before action. Any pipeline that lets a model's output reach a decision without independent verification carries a version of this risk, usually with lower stakes.