OpenAI reports 16,000 extraction attempts, names Moonshot associates, and publishes no attribution evidence
OpenAI says it disrupted a July campaign to extract the hidden reasoning of its models and attributes a core cluster to individuals associated with Moonshot AI. It gives dates and counts, says no encryption was broken and no database compromised, and publishes no technical evidence for the name.
By Parminder Kumar Sharma · · 19 min read

Sixteen thousand attempts, under four a user on average, and a name without evidence
OpenAI's post of 30 September says that on 24 and 25 July it saw 16,000 requests using an extraction pattern, from over 4,000 users. A footnote adds that these are "attempted, not necessarily successful, extractions". Divide one figure by the other and the average user made at most four requests across the two days. That is our arithmetic, and because both OpenAI figures are floors ("16,000" and "over 4,000"), four is a ceiling on the average, not a measurement. It is still a wide and shallow shape: whatever made this campaign visible, it was probably the pattern across accounts and not the volume from any one of them.
What the number does not establish is most of what the headlines hang on it. It is not a count of how much hidden reasoning anyone captured: OpenAI gives no total for the whole campaign, only for the two-day spike, and no figure for successes. It is not all attributed to Moonshot: OpenAI says it is "unclear whether all operators" were a single actor, and attributes only "a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi". The post names no person, gives no account count, publishes no indicator, and does not say that any model was trained on anything extracted. The attribution section is two sentences long.
OpenAI says it fully disrupted the campaign by 28 July and published on 30 September, 64 days later. This is a different story from our two earlier briefings on Chinese labs and Anthropic's Claude, on the Chinese regulator probe and on Xiaomi's open-weight model. Those rested on Anthropic's telemetry; this one rests on OpenAI's, and the method question is the same: the only party that can see the traffic is also the accuser. Neither earlier briefing is restated here.
The record, in dates
The dates do much of the work in this story, so here they are in one place, then drawn to scale. OpenAI's dates come from its own post. The rest come from the source named in the right-hand column.
Dates in and around the campaign, 2026, each with its source
| Date | What happened | Source |
|---|---|---|
| 23 Apr | White House memo NSTM-4 says the US has information that foreign entities, principally in China, run industrial-scale distillation. It names no company. | White House, NSTM-4 |
| 1 Jul | OpenAI says the campaign begins, at low volume. | OpenAI post |
| 16 Jul | Moonshot launches its Kimi K3 model. | The Register |
| 20 Jul | Researchers report the replay weakness to OpenAI, on their own timeline. | Researchers' update |
| 21 Jul | A Moonshot business lead says K3's gains come from original architecture, not a distilled copy. | Nandu Daily, via Secrss |
| 22 Jul | A White House official says Moonshot distilled Anthropic's Fable for K3. | Business Insider via AOL; The Register |
| 24 to 25 Jul | The spike: 16,000 attempted requests from over 4,000 users. | OpenAI post |
| 27 Jul | K3's weights first appear on Hugging Face, per the repository's first commit. | Hugging Face |
| 28 Jul | OpenAI says the campaign is fully disrupted. | OpenAI post |
| 10 Aug | The researchers' paper appears on arXiv. | arXiv 2608.09867 |
| 8 Sep | NSA, CISA and FBI advisory names Moonshot among six China-based firms. | CISA, AA26-251A |
| 10 Sep | Anthropic's report makes its own Moonshot allegations. | Anthropic, covered in our earlier briefing |
| 30 Sep | OpenAI publishes its account. | OpenAI post |
Three derived points follow. The campaign ran 27 days from start to disruption. The spike began 23 days in, and OpenAI says it was fully disrupted 3 days after the last spike day. And the spike came 8 days after Moonshot launched Kimi K3, which settles one thing by simple ordering: output collected on 24 and 25 July cannot be in the model Moonshot launched on 16 July. Moonshot's Hugging Face repository for K3 shows its first commit on 27 July, three days after the spike began, and no source says the posted weights differ from the launched model. Later models are a separate question, and no source says any was affected. OpenAI describes the weeks before the spike as low volume, and 15 of the 27 days fall before K3's launch.
The spike also followed the White House post about Moonshot by two days. I draw no inference from that; nothing in the sources links the two.
What is stated, and what is not
OpenAI's post of 30 September 2026 read against the questions a risk owner asks, with other sources where they speak
| Question | Stated | Not stated |
|---|---|---|
| Who | A core cluster is attributed to individuals associated with Moonshot AI. It is unclear whether all operators were one actor. | Any name, account count or company role. What "associated with" means: employees, contractors, resellers. |
| How it was attributed | It is called OpenAI's "assessment of attribution". | Any indicator, infrastructure link, payment link or timing analysis, or a reason for withholding them. The Hacker News notes no technical evidence was cited. |
| How big | 16,000 attempts from over 4,000 users on 24 and 25 July. Related prompt-pattern activity across more than 15,000 users. | A whole-campaign total, how many attempts worked, how many reasoning traces were captured, which models were targeted (The Register asked and had no reply). |
| Whether it was used | The activity is "consistent with adversarial distillation". | That any extracted reasoning trained any model, or that any Moonshot model changed. |
| Effect on users | No encryption broken, no database compromised, no direct access to stored conversations. A pathway letting someone who held another user's encrypted reasoning replay it was closed. | Whether any real user's reasoning was replayed or exposed, or whether anyone was notified. |
| Enforcement | Fraudulent accounts banned or restricted, signup and infrastructure controls tightened, third-party services worked with. | How many accounts, which third parties, any legal step. |
| Moonshot's reply | Nothing in the post. | Any reply. CNBC and The Register report no immediate response. The 21 July denial was about K3 in general and came first. |
| Government | Findings shared through the Frontier Model Forum and "appropriate government information-sharing channels". | Any government statement about this campaign. The 8 September advisory lists OpenAI models Moonshot is said to have distilled but cites no per-company evidence in the text I read. |
A note on the user counts. If the 4,000 sit inside the 15,000, which the post implies but does not say, they are about 27 per cent of it on the rounded figures. Both numbers are floors ("over 4,000", "more than 15,000"), so the share is indicative only.
One third-party measurement that I found bears on Kimi K3, and its authors disclaim the conclusion it invites. The researchers behind the paper OpenAI credits asked whether Kimi K3 responds unusually to reasoning decoded from Claude Opus 4.8 and GPT-5.6 Sol. Seeding K3's reasoning with a few words from those traces moved its style toward the source, more than it did for a control model; GLM-5.2, from a different lab, moved toward the Claude sample too.
Five comforting labels, and what each leaves out
Labels in OpenAI's post and the coverage, and the question each one skips
| Label | What it tells you | What it leaves out |
|---|---|---|
| Distillation | A training technique. Legitimate when an owner shrinks its own model; OpenAI qualifies this one as adversarial. | Authorisation. Whether it is allowed turns on the provider's terms, not on the technique. An OpenAI policy lead and the White House memo both say legitimate distillation is fine. |
| Protected reasoning | The model's internal working, withheld from the answer. | What protects it. In OpenAI's documented stateless mode it is an encrypted item that travels through your own application. The researchers say such blocks can never be more than semi-hidden. |
| Disrupted | OpenAI stopped the campaign by 28 July. | What was lost before it stopped. Disruption ends new attempts. It does not recall anything already captured, and the post gives no capture figure. |
| Individuals associated with Moonshot AI | Some people OpenAI links to the company. | The link. Employees, contractors, resellers and customers are different things. A core cluster is not the whole campaign. |
| Did not break our encryption | No intrusion into OpenAI's systems, on OpenAI's account. | Confidentiality. The weakness described is that an intact block could be replayed, and that the model itself reads it. No cipher had to fail. |
How to classify the campaign on the facts OpenAI states, and the basis for each judgement
| Frame | Fit on the stated facts | Basis |
|---|---|---|
| Intrusion or breach of OpenAI's systems | Not on OpenAI's account | Operators did not break encryption, compromise a database or reach stored conversations. |
| Personal data breach | Open, not stated | No exposure or notification is stated. The closed pathway concerned replaying another user's encrypted reasoning, and the researchers found personal data and credentials in such blocks published in public logs. Whether any real user's block was replayed here is not stated. |
| Intellectual property and contract | Fits | OpenAI frames it as a terms violation. Its Services Agreement bars using Output to develop competing models and defines reverse engineering to include "model extraction or stealing attacks". Whether any statute applies is a legal question the post does not raise. |
| Model security (adversarial machine learning) | Fits | The NCSC's taxonomy lists unauthorised distillation and model stealing under model inversion, and keeps ML-specific attacks apart from infrastructure compromise. |
| Supply chain | A question, not a finding | A model trained on extracted output carries the dispute to whoever deploys it. No source says any Moonshot model contains OpenAI-derived data. |
What was being copied, and where each control sits
OpenAI describes the technique in one sentence: operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. To see why that was possible, start with OpenAI's own documentation. In stateless mode, which applies when store is false or an organisation uses Zero Data Retention, reasoning items arrive with an encrypted payload that the client passes back on later turns. The reasoning is not shown to you, but it does pass through your systems.
The researchers OpenAI credits, in a paper of 10 August, found that such blocks were interchangeable across sessions, users and models inside a provider's family. On their account a weaker, less guarded model could be made to read out what a stronger one had reasoned, so the stronger model's refusals were never engaged. They disclosed to the providers first, and say their main results were no longer reproducible by August after provider mitigations. OpenAI says it confirmed the attack paths they identified were real. The Hacker News's coverage of the paper noted that it documents no malicious use in the wild.
How the operators in July came to use the same class of technique is not stated. The paper credits an earlier report of cross-session replay in May 2026, and the update's timeline marks the researchers' own "initial breach" of OpenAI's API on 18 July, after OpenAI says the activity began. Anthropic's September report describes a similar replay of Claude's reasoning signatures, attributed to Moonshot and DeepSeek, which our briefing on the regulator probe covers. Because the technique is now public, a resemblance between campaigns is weak evidence of who ran them.
Cost is not what stops this. The paper prices decoding 10,000 traces at about $720 on Claude Haiku 4.5 rates. Applied to 16,000 requests that gives about $1,150. This is arithmetic on a different provider's price list, not an estimate of anyone's spend. It makes one point: at this scale an extractor needs a way past the controls, not a budget.
Read the diagram as a defender. A per-account limit, at point 2, sees little when the average is under four requests a user. The weight falls on pattern checks across accounts (3), on binding the item to its context (4), on holding streamed reasoning (5) and on parity across hosts (6). OpenAI says it has acted on 1, 3, 4 and 5 in some form, and says the work on 6 is not finished.
The researchers' September update, listed in the sources, says the fixes were uneven. Their tests on 13 September found the replay route blocked on OpenAI's own API but still working intermittently on OpenAI models served through Microsoft Azure, and a second route that involves no decryption still returning reasoning from OpenAI models. These are the researchers' claims; I have not tested them and read no vendor confirmation. Their own timeline records Azure mitigations in the last days of September. They line up with OpenAI's closing admission that partner-hosted deployments "need the same protections as first-party services" and that tool-output attacks need protections that look beyond visible text.
What a provider should do, and what a buyer should ask
If you offer a model API, the sourced UK baseline is the voluntary Code of Practice for the Cyber Security of AI (31 January 2025). Provision 6.2 says a developer offering an API "shall apply controls that mitigate attacks on the AI system via the API", with limits on access rate as the example. Provision 9.4 asks developers to evaluate outputs so they do not let users reverse engineer non-public aspects of the model. Provisions 12.1 and 12.2 ask system operators to log activity and analyse it for anomalies. The US advisory AA26-251A adds behavioural indicators: usage with no human idle periods, new accounts at maximum usage straight away, subscription-to-usage ratios that do not fit, and stronger identity verification on accounts.
The advisory also recommends something buyers should notice. It tells providers to "subtly alter responses" for suspected distillation attempts, and not to tell suspected distillers that they have been switched to a downgraded model.
The design choice underneath is where the reasoning lives. The researchers list the options: keep it on the server and hand the client only an identifier; bind each block to its user, conversation and model; stop models reading each other's blocks at the gateway; revoke blocks on anomalous replay; or drop the encryption for older models. Each has a cost. Anthropic's documented example binds a thinking block to the exact prefix that produced it, so integrators must keep requests append-only, and that rule is enforced by default only for accounts created from 31 August 2026. Tight binding can break proxies, gateways and tools that rewrite conversation history, so part of the cost lands on customers. OpenAI has not described its own fix in that detail.
For a buyer, four things follow from the sources.
First, you may be holding the artefact. In stateless mode, including Zero Data Retention, the encrypted item travels through your application. The researchers decoded 315,320 blocks taken from public agent logs and recovered 367 personal-data artefacts and 182 credentials, and in some cases the recovered data had not appeared in the user's own input. On OpenAI's documentation, a privacy control that keeps data off the vendor's servers moves a sensitive artefact into your logs, and sanitising the visible text does not reach it. That is the researchers' finding about published logs, not an OpenAI incident.
Second, the route matters. First-party and cloud-partner endpoints can differ in what protections they carry and when. Ask for the same-protections commitment OpenAI says partner deployments need, in writing, with dates.
Third, silent degradation. A provider acting on the advisory may quietly give a suspected extractor worse answers. That is reasonable against an attacker and a risk for a customer caught in a false positive. OpenAI's 15,000 users are a cluster of prompt-pattern activity, and the post does not say how many were banned, restricted or cleared.
Fourth, your own distillation rights. OpenAI's Services Agreement (effective 1 January 2026) bars using Output to develop competing models, with a Permitted Exception for classifiers and embeddings that are not distributed to third parties and for fine-tuning through OpenAI's own services. The Terms of Use for individuals, which cover UK residents, carry the same bar in one line.
The UK position, as far as it is sourced
None of the sources says anything about UK users, or about UK law applying to this campaign. OpenAI's post names no jurisdiction beyond "government information-sharing channels". What can be sourced is the framework a UK security lead would use.
On contract, OpenAI's Services Agreement names OpenAI OpCo, LLC as the contracting party for customers outside the EEA and Switzerland, and sets Irish law as the governing law for customers in the EEA, Switzerland and the UK. UK individuals contract with OpenAI OpCo, LLC under the Europe Terms of Use, updated 16 January 2026.
On technical classification, the NCSC's paper on adversarial attacks against machine learning (published 29 April 2026, last modified 12 August 2026) lists unauthorised distillation and model stealing as example techniques and separates ML-specific attacks from infrastructure compromise. That is the same line OpenAI draws, and it is why this belongs in a model-security register and not an incident log for OpenAI's systems.
On data protection, the Code of Practice says that where data used for an AI system is personal, those involved may have data protection obligations and should consult ICO guidance. Whether any UK person's data was in an extracted block here is not stated.
What to do, in the order worth doing it
Take this with you
Checks for UK providers and buyers of model services
- Find every place your applications store, log or forward model reasoning artefacts: encrypted reasoning items, thinking blocks and signatures. Start with stateless and Zero Data Retention deployments, agent frameworks and observability tools.
- Strip those artefacts, as well as the visible text, from any agent log, benchmark run, bug report or dataset before it leaves the organisation or goes into a repository, and add a pipeline check that blocks them.
- Record which route each model reaches you by: the vendor's own API, a cloud partner or an aggregator. Ask the vendor in writing whether the partner route carries the same reasoning protections as the first-party API, and since when.
- Check how staff and systems obtain API keys. Buying, selling or transferring keys breaches OpenAI's terms, and resold access is one of the routes the US advisory describes. Retire any key you cannot trace to a contract.
- If you offer a model API or a model-backed service, measure what your own limits can see: compute average requests per account for your busiest hour, then add checks that work across accounts, such as signup attributes, payment methods and timing, and log enough to correlate.
- If you plan to train or distil a smaller model from a vendor's outputs, read the vendor's output-use clause first and record the decision. OpenAI's Services Agreement allows only its Permitted Exception.
- For each vendor model in production, ask who trained it on what, and record the answer or the silence, using the provenance questions in our earlier briefing on Xiaomi's open-weight model. Record the Moonshot attribution as reported and unproven, with a review date, not as a fact about any Kimi model. Ask what the contract says if a vendor's model is later shown to have been trained on outputs obtained in breach of another provider's terms: warranty, indemnity and a right to terminate.
- Ask each AI vendor how it treats accounts it suspects of extraction: whether responses are degraded without notice, what notice and appeal you would get if your account were swept up, and how it would tell you.
- Set review triggers: a Moonshot reply, any published indicator or technical evidence, a regulator or court step, and any vendor change to how reasoning items are bound.
Method, motive and the counterargument
OpenAI has a direct commercial interest. Moonshot competes with it, and the post argues for a national security risk and for coordination with government. The Register's reply is blunt: its headline calls the complaint ironic, and its standfirst says US model makers can train on web data but call distilling theirs a national security risk. That is an argument about consistency, and a fair one to put to any company that collected the web and now polices outputs. It does not touch the two questions this briefing turns on: whether the activity described happened, and whether a design that lets a hidden-reasoning block be replayed is a security weakness. Independent researchers found the weakness across three providers, not only OpenAI.
The counterargument has a political form too. Bloomberg, as relayed by The Next Web, reports David Sacks calling such reports a push to ban rival open models, and Caroline Zier, who leads national security policy initiatives at OpenAI, replying that the concern is "violation of our terms of service, not open models or legitimate distillation". I could not read the Bloomberg piece, which sits behind a bot check I did not bypass, so both statements are second-hand. The White House memo of 23 April also says legitimate distillation is fine, then calls industrial campaigns "unacceptable". None of the sources I read tests whether the terms OpenAI relies on would hold up; they are contract terms a vendor writes for itself.
Keep method and accusation apart. The method, replaying a hidden-reasoning block across conversations, is documented by independent researchers, confirmed as real by OpenAI and described again by Anthropic. The accusation, that Moonshot-linked people ran it and for what, is OpenAI's assessment alone, with no published indicator. OpenAI also competes with Anthropic, whose documentation and report I cite here for comparison.
The question that exposes the gap
OpenAI says related prompt-pattern activity ran across more than 15,000 users and that it banned or restricted fraudulent accounts. It does not say how many of the 15,000 were fraudulent, whether any were ordinary customers, or how anyone would learn which they were.
So the question for your own risk register is this: if one of those 15,000 accounts belonged to your organisation, or to a router your staff use, what in your logs, or in your vendor's notice terms, would have told you before 30 September?
Key facts
Sources
- PrimaryDisrupting a coordinated model-distillation campaign, 30 September 2026. Read in full in a browser (curl returned 403). Source of every OpenAI figure, quote, date and mitigation in the briefingOpenAIaccessed 2026-10-01
- PrimaryStealing Reasoning Traces from Proprietary LLM APIs, 10 August 2026. Read in full: mechanism, disclosure, mitigations, the privacy results and the Kimi K3 probe in Appendix BPanfilov, Schmotz, Shumailov and colleagues (arXiv)accessed 2026-10-01
- PrimarySeptember 2026 update on whether the fixes held, with a status table dated 13 September and a July to September timeline. Used for the claimed gaps on cloud partners, which are the authors' claims and unconfirmedThe paper's authorsaccessed 2026-10-01
- PrimaryJoint advisory AA26-251A of 8 September 2026 on China-based distillation campaigns. Read in full for what it says about Moonshot, the OpenAI models it lists, the indicators and the recommended mitigationsCISA, NSA and FBIaccessed 2026-10-01
- PrimaryNSTM-4, Adversarial Distillation of American AI Models, 23 April 2026. A two-page scanned memo read from the page imagesWhite House Office of Science and Technology Policyaccessed 2026-10-01
- PrimaryOpenAI Services Agreement, effective 1 January 2026. Read for section 3.3 restrictions, the Permitted Exception, the contracting party and governing law for UK customersOpenAIaccessed 2026-10-01
- PrimaryEurope Terms of Use, updated 16 January 2026, which cover UK residents. Read for the Output and extraction restrictions and the contracting partyOpenAIaccessed 2026-10-01
- PrimaryReasoning models guide in the API documentation. Used for stateless mode, Zero Data Retention and the encrypted reasoning item the client passes backOpenAIaccessed 2026-10-01
- PrimaryPreserved thinking documentation. Used as a worked example of binding a reasoning block to its context, and for the integration cost and the 31 August 2026 enforcement dateAnthropicaccessed 2026-10-01
- PrimaryDetecting and countering misuse of AI, September 2026. Pages 146 to 150 only, to confirm the cross-session replay technique it attributes to Moonshot and DeepSeek. Not restated; see the earlier briefingAnthropicaccessed 2026-10-01
- PrimaryUnderstanding adversarial attacks against Machine Learning and AI, published 29 April 2026 and modified 12 August 2026. Used for the classification of model stealing and the line between ML attacks and infrastructure compromiseNCSCaccessed 2026-10-01
- PrimaryCode of Practice for the Cyber Security of AI, 31 January 2025, voluntary. Read for provisions 6.2, 9.4, 12.1 and 12.2 and the data protection noteUK Governmentaccessed 2026-10-01
- PrimaryKimi K3 model repository. Its public commit history (Initial commit, 27 July 2026, 13:31 UTC) dates the first appearance of the weights; the page also names the custom Kimi K3 LicenseHugging Face (Moonshot AI repository)accessed 2026-10-01
- Reported byThe 1 October report that prompted this briefing. Used as a pointer to the OpenAI post and for the note that no technical evidence was citedThe Hacker Newsaccessed 2026-10-01
- Reported byThe 30 September sceptical piece. Read for the counterargument, the unanswered question about which models were targeted and the context on the White House claimThe Registeraccessed 2026-10-01
- Reported by23 July report on the White House accusation against Moonshot, and the source for the 16 July release of Kimi K3The Registeraccessed 2026-10-01
- Reported by22 July 2026 report on the White House official's post on X about Moonshot and Anthropic's Fable. Used to date the post to 22 JulyBusiness Insider via AOLaccessed 2026-10-01
- Reported byReport that Moonshot did not immediately respond to a request for comment on the claimCNBCaccessed 2026-10-01
- Reported byReport relaying Bloomberg for the OpenAI spokesperson's statement and David Sacks's criticism. Bloomberg itself was behind a bot check and was not readThe Next Webaccessed 2026-10-01
- Reported byChinese-language report of Moonshot's 21 July response to the distillation doubts about Kimi K3, which also dates the K3 release to 16 JulyNandu Daily, via Secrssaccessed 2026-10-01
- Reported by12 August coverage of the researchers' paper, used for the note that it documents no malicious exploitation in the wildThe Hacker Newsaccessed 2026-10-01


