An AI written intelligence report reached airborne aircraft before anyone opened its sources
Four anonymous sources told CNN that a chatbot generated intelligence claiming a Chinese ship carried nuclear weapons components, and that the operation was stopped only at the last moment. The control that failed was provenance, not the human in the loop.
By Parminder Kumar Sharma · · 18 min read

The whole record is four people who will not be named
Count the on the record confirmations behind the defence story of the week: none. CNN's account rests on four sources familiar with the episode, not one of them named. US Special Operations Command Pacific and the Pentagon were both asked to comment and neither responded. No document has been released. The chatbot is not named, and CNN says it could not establish whether it was a commercial product or a government one. CNN also reports that it could not learn what the ship was actually carrying.
The one figure in this story you can check for yourself comes from somewhere else entirely. The Department of War's generative AI platform, GenAI.mil, launched on 9 December 2025. On 12 June 2026 the department's chief technology officer, Emil Michael, said that 1.5 million people out of three and a half million were using AI every day, against 80,000 users in December. That is 185 days and a factor of 18.75, and it takes the platform to roughly 43 per cent of the workforce.
What that number does not establish is that GenAI.mil had anything to do with this episode. Google described the platform as authorised at Impact Level 5 for controlled unclassified information and intended for unclassified work. The episode as described involved secret signals intelligence held in government systems, so the platform most readers have heard of is not, on its own, a candidate. The January 2026 strategy does say the aim is to put models in front of personnel at all classification levels, which is where a classified equivalent would live, but CNN does not name a system and this briefing will not guess at one.
Everything that follows about the episode itself carries the same qualifier: it is the account of four people who spoke anonymously to two CNN reporters, and no part of it has been confirmed on the record by any department.
What CNN's sources state and what the account leaves open (CNN, 18 September 2026)
| Stated | Not stated |
|---|---|
| A report claimed a Chinese ship in the Middle East carried components of a nuclear weapons programme | What the cargo actually was. CNN reports it could not learn |
| A chatbot fused open source intelligence with secret signals intelligence in government holdings | Which chatbot, and whether it was commercial or a government product |
| The analyst used AI again to package the findings into a standard intelligence report and disseminated it | Whether the product recorded anywhere that AI had been used |
| Armed personnel were preparing to board and military planes were in the air | The date, the vessel, the location beyond the Middle East |
| Officials dug deeper into the report just before the planned operation | Who caught it, what prompted the second look, and who stopped the operation |
| One source called the report entirely false | Any correction, review, inquiry, discipline or policy change since |
The chain, step by step
The sequence CNN describes is short. An analyst at a special operations command had intelligence reporting on a Chinese ship's manifest, reporting that originated with US Special Operations Command Pacific in Hawaii. The analyst put a question to a chatbot. The tool drew on open source material and on secret signals intelligence held in government systems, and concluded that the ship was carrying material associated with a nuclear weapons programme. That conclusion was wrong.
The analyst then used AI a second time, to turn the findings into a standard intelligence report, and sent it out. CNN describes the format precisely: it is the kind of product "that is trusted by military officials". The report circulated across the US military, interception planning began, armed personnel prepared to board and aircraft were airborne. Only at that point did officials dig into the underlying report, find that it had been generated with the help of AI, and find the identification false.
Read that as a pipeline rather than as a drama and two things stand out. The first is that the failure is not exotic. A model produced a confident wrong answer about what was in a container, which is the most ordinary failure mode these systems have. The second is that nothing in the chain, as described, carried the fact of the model's involvement forward with the product. The information that would have let any reader downgrade the report travelled separately from the report, in the heads of the people who made it.
The second AI pass is where the integrity went
The first use of AI produced a wrong answer. Wrong answers are survivable when the product says where it came from. The second use is the one that did the damage, because it took a chatbot answer and put it inside a container whose whole function is to signal that a claim has been through the process that normally produces such claims.
That is an integrity failure in the strict sense. Nothing was stolen, no adversary was involved, no system was breached. The content simply was not what its container asserted about it. A reader three commands away had no way, from the artefact in front of them, to tell the difference between a judgement built by an analyst against described sources and a judgement produced by a language model and then formatted to look like one.
The intelligence community has had a written answer to exactly this problem since 2 January 2015. Intelligence Community Directive 203 sets out five analytic standards and nine analytic tradecraft standards. The first tradecraft standard is that analytic products "should identify underlying sources and methodologies upon which judgments are based", using source descriptors that describe factors affecting source quality and credibility. The third is that products "should clearly distinguish statements that convey underlying intelligence information used in analysis from statements that convey assumptions or judgments".
Two caveats, stated rather than buried. ICD 203 applies to the Intelligence Community as defined in statute, and CNN does not establish that this product was a formal all source analytic product subject to it, nor whether it carried a source summary. And ICD 203 predates generative AI by years: it says nothing about a model as a producer of text, because in 2015 there was nothing to say.
The human in the loop was never missing
It is worth being blunt about the most comforting phrase in this field. There was a human in the loop at every step of the chain CNN describes. A human wrote the prompt. A human decided the answer was worth disseminating. Humans read the report, humans planned the interception, humans got into the aircraft. The loop was full of people and it did not help, because none of them had the one thing that would have let them object: sight of how the claim was produced.
One of CNN's sources puts the same point in operational terms, saying of targeting that "there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide". A human in the loop is an arrangement of authority. It is not, by itself, a verification control, and calling it one is how organisations end up believing they have a gate when they have a signature.
Comforting labels against what they actually constrain, from the documents and remarks cited in this briefing
| Label | What it constrains | What it does not touch |
|---|---|---|
| Human in the loop | Who holds authority for the decision | Whether that person can see how the input was produced |
| Guardrails | What a user is permitted to ask the tool to do | Whether the answer is true or its sourcing is described |
| A standard intelligence report | That the document looks like others of its class | That it was produced the way that class is normally produced |
| Responsible AI, as redefined in the January 2026 strategy | Freedom from ideological tuning, per law firm analyses of the memo | Verification of model output before it is disseminated |
On guardrails, the department's own framing is on the record and is worth quoting against this episode. Speaking on 12 June, Emil Michael said there is "risk in everything" but that guardrails are in place, adding: "Right now at this level we are not talking about the warfighters in Central Command, we are talking about the average employee." He also said target lists are still developed "independent of AI". Neither statement is dishonest, and neither covers the case CNN describes, which is an intelligence product moving through a warfighting decision chain in the spring, before those June remarks were made.
The rules that already existed, and what they do not do
The United States is not short of written commitments here. On 9 November 2023 it published the Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy, which 58 states had endorsed as of 27 November 2024, the United Kingdom among them. Two of its measures are directly on point. States should ensure military AI capabilities are developed with "methodologies, data sources, design procedures, and documentation that are transparent to and auditable by their relevant defense personnel". And personnel who use or approve the use of such capabilities should be trained to understand their capabilities and limitations "to mitigate the risk of automation bias".
What the declaration does not do is bind anyone. It is political, not legal, and it contains no mechanism that stops an unmarked product moving through a distribution list.
Three documents read in full for this briefing, and the gap each leaves
| Document | What it requires | What it does not do |
|---|---|---|
| ICD 203, Analytic Standards, 2 January 2015 | Products identify underlying sources and methodologies, and separate intelligence from assumption and judgement | Name a model as a source type, or require any marking of AI use |
| Political Declaration, 9 November 2023, 58 endorsing states | Auditable methodologies and documentation, and training against automation bias | Bind any signatory, or create a gate in any pipeline |
| Artificial Intelligence Strategy for the Department of War, 9 January 2026 | An AI first force, with models at all classification levels and data made available for exploitation | Set a common verification standard, which CNN's sources say does not exist |
The adoption curve this sits on
The context is not a cautious pilot. The 9 January 2026 memorandum directs the department to become an "AI-first" warfighting force "across all components, from front to back". The strategy is built on three pillars, warfighting, intelligence and enterprise operations, executed through seven pace setting projects, one of which is described as accelerating the conversion of intelligence into weapons. In the same period the department went from 80,000 AI users to 1.5 million daily users, created more than 100,000 customised agents, and had 50,000 people signed up for training with a waiting list.
The use case the department is proudest of is worth looking at closely, because it has exactly the shape of the failure. Michael described loading papers into the system to have it draft a report to Congress that would otherwise take 200 hours of staff time, in five hours. Load the source material, have the model produce the product, send the product to people who will act on it. That is the congressional reporting workflow and it is also, in CNN's account, the intelligence reporting workflow. The difference is not the pipeline. The difference is what happens at the other end.
There is a commercial dimension worth naming without sneering. The platforms in question are commercial models sold into government at speed, and the vendors have an obvious interest in adoption figures being quoted as evidence of success. A former senior US official told CNN that "the internal tools are mostly just copies of the commercial stuff wearing lipstick". Adoption is a sales metric. It is not an assurance metric, and nobody in the reporting claims it is.
Why this belongs in a security briefing at all
A hallucination is not a security incident in the way the word is normally used. There is no threat actor, no unauthorised access, no exfiltration, no confidentiality loss, nothing that would trigger a notification to the Information Commissioner or anyone else. If your incident taxonomy only has room for confidentiality, this event does not appear in it at all.
It is, however, an integrity failure inside a decision pipeline, and integrity is the leg of the triad that enterprises are quietly removing as they let agents draft the documents that people act on. The transferable control is not a model evaluation and it is not a policy about which chatbot staff may use. It is the pair of gates that keep a claim attached to its evidence: a marking that says how the artefact was produced, and an independent check of the underlying records before anything irreversible happens.
The difference between the two cases should be stated plainly rather than smoothed over. In the military case, a verification step did exist in doctrine and a human did eventually perform it, late, and the operation did not happen. Nobody boarded anything, nobody was hurt, and the piece of this story that actually worked was a person deciding to read the underlying report. In most enterprises deploying agents to write reports, there is no equivalent doctrinal step at all, no convention for marking which paragraphs a model produced, and no habit of opening the source records before acting. The consequences are also different by orders of magnitude, and flattening a near miss involving a Chinese vessel into a lesson about quarterly risk reports would be silly. What transfers is the shape of the failure, not the stakes.
What a UK reader can actually check
Nothing in the reporting touches UK systems, and this briefing makes no claim about them. What can be checked is what the UK has written down, and it is more specific on this exact point than anything found on the US side.
JSP 936 Part 1, Dependable Artificial Intelligence in Defence, version 1.1 of November 2024, describes the failure mode before it describes the controls. At paragraph 37 it warns that "credibly incorrect outputs from generative AI (including Large Language Models, LLMs) may mislead human operators into poor decision-making". Paragraph 39 requires that digital systems containing AI in scope of the JSP are clearly identified as AI based. Paragraph 40 is the one that matters most here: products such as documents that have had AI applied in their development "must clearly state that AI has been used in their production", and for high risk products, information on how the AI was developed, used and assured "should be available to the risk owner of decisions being made on the basis of the product". Paragraph 42 adds that where AI provides advice to an operator, alternate sources and their provenance should be identified.
Read against CNN's account, paragraph 40 is the missing gate stated in a single sentence: the product must say that AI made it, and the person carrying the risk must be able to see how it was assured. The MOD's own cover note is candid that the JSP represents an aiming point, with supporting tools and frameworks still in development, so a UK reader should treat it as published policy rather than as evidence of practice. It is worth noting that the JSP itself discloses that an AI and human partnership supported the production of its cover page and foreword, and that the rest was written and reviewed by humans. That is the marking rule applied to the document that contains it.
On the civil side, the DSIT AI Cyber Security Code of Practice of 31 January 2025 carries the same logic in voluntary form. Principle 4, enable human responsibility for AI systems, says at 4.3 that where human oversight is a risk control, developers or system operators "shall design, develop, verify and maintain technical measures to reduce the risk through such oversight". Principle 8 requires documentation of data, models and prompts. Principle 12 requires logging of system and user actions and analysis of those logs. The UK also endorsed the 2023 Political Declaration, so the automation bias and auditability measures are commitments it has made in its own name.
Verification gates for machine written analysis, in the order worth doing them
Take this with you
Do these in order
- List the products people act on without re-deriving the facts: incident reports, risk assessments, due diligence summaries, threat intelligence, board papers, supplier assurance packs.
- For each one, write down what an irreversible action looks like: money moves, a contract is signed, a system is isolated, a person is suspended, a customer is notified.
- Require every such product to carry a marking of whether a model drafted or formatted it, and which parts. Treat an unmarked product as unverified rather than as human written.
- Require a source line for each material claim, naming the underlying record rather than the model or the summary that produced it.
- Separate generation from formatting. Never let one pipeline both reach a conclusion and dress it in the template that signals the conclusion was checked.
- Put one named person between the product and any irreversible action, and give that person both the authority to stop and the time to open the underlying records.
- Log prompts, retrieved documents and outputs so a later reader can reconstruct how a claim was produced, and keep those logs as long as the decisions they support.
- Test the gate rather than the model. Sample finished products and try to trace each material claim back to a record. Count the ones you cannot.
- Train explicitly for automation bias, and measure whether reviewers actually open sources rather than whether they attended the session.
- Set one rule for high consequence decisions: no single machine written product initiates an irreversible step on its own.
The order matters. Marking before verification is pointless, because there is nothing to act on. Verification before you have listed the irreversible actions is expensive, because you end up checking everything at the same depth. The cheap version of this control is narrow and specific: find the handful of products that move money, people or systems, and make those carry their provenance.
The question that exposes the gap
Jake Steckler of GovAI, a former US Army officer, told TechCrunch that it is especially critical for decisions that could lead to use of force that service members understand the uncertainty inherent to large language models, and that prioritising adoption speed above all else will produce incidents that cost trust. That is the right frame, and it applies well outside defence.
So ask your own organisation the question this episode asks of the Department of War, and do not accept the comfortable version of it. The comfortable version is: do we have a human in the loop. The exposing version is: if a report written by a model were wrong tomorrow, what would the person approving the action have to open in order to disagree with it, and do they have it. In the account CNN's four sources give, the answer was: the underlying report, and yes, eventually, with aircraft already in the air.
Sources
- PrimaryIntelligence Community Directive 203, Analytic Standards, 2 January 2015: the five analytic standards and nine tradecraft standards, including sourcing and the separation of intelligence from judgementOffice of the Director of National Intelligenceaccessed 2026-09-19
- PrimaryJSP 936 Part 1 Directive, Dependable Artificial Intelligence in Defence, V1.1 November 2024: paragraphs 37, 39, 40 and 42 on misleading generative AI outputs and marking AI use in productsUK Ministry of Defenceaccessed 2026-09-19
- PrimaryAI Cyber Security Code of Practice, 31 January 2025: principles 4, 8 and 12 on human responsibility, documentation and monitoringDepartment for Science, Innovation and Technologyaccessed 2026-09-19
- PrimaryPolitical Declaration on Responsible Military Use of Artificial Intelligence and Autonomy, 9 November 2023: the measures on auditable methodologies, automation bias training and senior oversightUS Department of Stateaccessed 2026-09-19
- PrimaryLanding page listing the 58 endorsing states as of 27 November 2024, including the United Kingdom and the United StatesUS Department of Stateaccessed 2026-09-19
- PrimaryPress release, 9 December 2025: Gemini for Government as the first AI on GenAI.mil, IL5 authorisation, three million personnel, unclassified workGoogle Cloudaccessed 2026-09-19
- PrimaryRelease note for the AI Acceleration Strategy: three pillars and seven pace setting projects, including accelerating the conversion of intelligence into weaponsUS Under Secretary of War for Research and Engineeringaccessed 2026-09-19
- Reported byKatie Bo Lillis and Zachary Cohen, 18 September 2026: the originating account of the episode, its anonymous sourcing, the two AI steps, the SOCPAC manifest reporting and the quotes used hereCNNaccessed 2026-09-19
- Reported byFollow up coverage of the CNN report, used for the GenAI.mil and Political Declaration context it addsArs Technicaaccessed 2026-09-19
- Reported byAditya Mehta, 18 September 2026: follow up coverage, used for the Jake Steckler comment on uncertainty in LLMsTechCrunchaccessed 2026-09-19
- Reported byMatthias Bastian, 19 September 2026: follow up coverage, checked for any detail beyond CNNThe Decoderaccessed 2026-09-19
- Reported by12 June 2026: Emil Michael on 1.5 million daily users against 80,000 in December, guardrails, training numbers, custom agents and target listsDefenseScoopaccessed 2026-09-19
- Reported byAnalysis of the 9 January 2026 Artificial Intelligence Strategy for the Department of War, used for the pace setting projects and the redefinition of responsible AI, since the strategy PDF could not be retrieved directlyCovington and Burling, Inside Government Contractsaccessed 2026-09-19
- Reported byKnowledge base entry quoting the 9 January 2026 memorandum directing an AI first warfighting forceDefense Management Instituteaccessed 2026-09-19


