P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

MiMo-V2.6-Pro leads the open-weight index, and the provenance question it leaves open

Xiaomi's MiMo-V2.6-Pro is the highest scoring open-weight model on the Artificial Analysis Intelligence Index, and Anthropic says Xiaomi replayed its own users' sessions through Claude. Neither party's published record answers the question a UK buyer has to ask.

By Parminder Kumar Sharma · · 15 min read

Editorial illustration for the briefing: MiMo-V2.6-Pro leads the open-weight index, and the provenance question it leaves open

Add up the exchanges, then look at the dates

Anthropic's September 2026 threat intelligence report puts a number on five of the distillation campaigns it attributes to Chinese labs: over 151 million exchanges for Alibaba, over 23 million for Moonshot, over 12.1 million for DeepSeek, over 3.4 million for Zhipu, and over 400,000 for Xiaomi. Those five add to 189.9 million. Xiaomi's share is 0.21 per cent of them, roughly one exchange in 475, and the smallest of the named campaigns by a factor of eight.

That 400,000 is the entire published basis for the line now running across the trade press: that Claude helped get Xiaomi's new model to the top of the open-weight rankings. Anthropic observed those exchanges over 20 days in March and April 2026. Xiaomi uploaded the MiMo-V2.6 weights to Hugging Face on 21 September 2026, at least 144 days after the end of April, and says the reinforcement learning run that produced them took under six days.

Here is what the arithmetic does not establish. It does not show that Anthropic is wrong, that the campaign was small in effect, or that 400,000 agentic coding sessions are a trivial training signal. Anthropic states in the same report that distillation can deliver significant uplift using fewer exchanges than the campaigns it describes. The arithmetic establishes something narrower and more useful to a buyer: the public record contains a serious, specific allegation about how Xiaomi collected training data five months before this model shipped, and it contains no statement at all, from either party, about what went into MiMo-V2.6-Pro.

That gap is the story. A UK organisation deciding whether to put an open-weight model into a regulated workflow does not need to referee a dispute between two model vendors. It needs a provenance statement, and on 22 September 2026 there was not one.

What Xiaomi actually shipped

The MiMo-V2.6 series is three models. MiMo-V2.6-Pro is a sparse mixture-of-experts model with 1.02 trillion total parameters and 42 billion active per token, a 1M token context window, and text, image, video and audio inputs. MiMo-V2.6-Flash is 309 billion total and 15 billion active. A third variant, Pro-UltraSpeed, is a faster serving mode of the same model at up to 20 times the output speed, priced ten times higher. The weights for Pro and Flash are on Hugging Face, along with a 9 billion parameter distilled model intended as a starting point for other people's training runs.

Xiaomi's own release post is specific about the training it wants to talk about. Each of Pro and Flash completed 30 reinforcement learning steps over roughly 750,000 trajectories in under six days, at about $2.62m for Pro and $0.85m for Flash. The stated batch shape, 1,568 prompts times 16 rollouts per step across 30 steps, works out at 752,640 trajectories, which matches the roughly 750,000 figure. On DeepSWE v1.1, a held-out long-horizon software engineering benchmark, Xiaomi reports Pro rising from 58.4 to 72.57 and Flash from 48.8 to 65.68.

What the release post, the model card and the collection page do not contain is any statement about where the data came from, at any stage. The card describes a technique it calls Multi-Prefix Multi-Teacher On-Policy Distillation, which reuses "histories from teacher trajectories", and never names a teacher. There is no pre-training data description, no data provenance section, and no response to the allegation that was published eleven days earlier.

Xiaomi's own published material for MiMo-V2.6, read on 22 September 2026: release post, Hugging Face model cards and collection page

Question a buyer asksStated by XiaomiNot stated by Xiaomi
Model size and shape1.02T total, 42B active, 1M context, MoE with 384 routed expertsNothing material: the architecture is documented in full
LicenceA metadata tag reading mit on each model cardNo LICENSE file in the repository, no copyright holder named, no terms for outputs
Reinforcement learning30 steps, about 750k trajectories, under six days, about $2.62m for ProNothing on the pre-training corpus that preceded it
DistillationMulti-teacher on-policy distillation after the RL phaseWhich teacher models, whose outputs, obtained how
BenchmarksA comparison table including Claude Opus 5 and GPT-5.6 SolWho ran the comparison numbers, and under which harness
The Anthropic allegationNo mentionNo denial, no correction, no acknowledgement

What the number 46 measures, and what it does not

The ranking claim is checkable, and it holds. Artificial Analysis scores MiMo-V2.6-Pro at 46.32 on version 4.3.2 of its Intelligence Index and places it first of 114 models in its comparison class of large open-weight models. The index is a weighted composite of ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.

The cost claim holds too, and it is the part that will move procurement. Artificial Analysis puts the weighted cost of one index task at $0.13 for MiMo-V2.6-Pro against $5.98 for Claude Opus 5.5, a ratio of about 46 to 1, on list prices of $0.435 and $0.87 per million input and output tokens against $4.00 and $20.00. Xiaomi kept V2.5 pricing for a model that scores 20.33 points higher than MiMo-V2.5-Pro did five months ago.

Three qualifications belong next to that headline. First, first among open weights is not first: on the same index, Claude Opus 5.5 scores 57.62, Claude Fable 5.1 scores 53.35, GPT-6 Astra scores 52.67 and Muse Spark 1.3 scores 48.09. MiMo-V2.6-Pro sits at 80 per cent of the leading score, and its lead over the next open-weight model, GLM-5.3 at 44.78, is 1.55 points.

Second, the index is measured against the vendor's own hosted API, not against the weights you would run yourself. Artificial Analysis states that its figures represent the model's first-party API, or the median across providers where there is no first-party API. A self-hosted deployment is a different artefact: the published Hugging Face repository stores the weights in fp8, about 566GB of files, and Xiaomi's own recipe serves Pro with tensor parallelism across 16 accelerators on two nodes. Nothing guarantees that your quantisation, your harness and your serving stack reproduce a score measured elsewhere.

Third, the benchmark table on the model card is not independent. Hugging Face's own eval-results block for the repository records four scores, including DeepSWE at 71.9, marked as unverified and added by a pull request rather than by a third-party run. Xiaomi's release post gives 72.57 for the same benchmark. The two numbers differ by 0.67 points, which is a small thing in itself and a useful reminder that a benchmark number without a stated harness is a marketing figure, not a measurement.

Anthropic's claim, in Anthropic's words

Anthropic's report, published in September 2026 and first covered on 11 September, covers activity it says it disrupted between December 2025 and August 2026. Its distillation chapter names seven labs: Alibaba, Moonshot, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax. Anthropic says it attributes the campaigns "with high confidence" to specific PRC-based labs, and calls the practice illicit distillation, distinguishing it from distillation as a legitimate training method.

The Xiaomi case is tagged GTG-16008. In Anthropic's account, Xiaomi replayed user conversations and coding sessions from its own MiMo models to Claude, often through the OpenClaw and OpenCode coding harnesses, and saved the exchanges to generate data for both supervised fine-tuning and reinforcement learning. Anthropic says it observed more than 400,000 requests across more than 1,500 accounts on proxy services, and that Claude was used to reconstruct developer environments from transcripts, to clean multi-turn conversations, to generate synthetic request and response pairs, and to judge answer quality.

Two details in that account matter more to a UK buyer than the headline does. Anthropic states that its investigation "did not indicate that Xiaomi used Claude's responses to serve its users", which is the opposite of what it alleges against Moonshot and DeepSeek. And it states that the relayed requests "contained the names, contact information, corporate data, and other sensitive data from hundreds of Xiaomi users in at least a dozen languages", much of it routed through third-party model routing services that Anthropic describes as commonly used by users in the United States and Europe. Anthropic's own characterisation is that these practices are "likely inconsistent with privacy laws and the labs' own terms of service".

A five step timeline beside a panel. March to April 2026, Anthropic observes over 400,000 requests it attributes to Xiaomi. September 2026, it publishes case GTG-16008 on replayed MiMo sessions used for future models, naming no checkpoint. 21 September, the weights appear on Hugging Face with an MIT tag and no training data statement. 22 September, Artificial Analysis scores 46.32, first among open weights. The panel separates what is on the record from what nobody has established.
Drawn from Anthropic's September 2026 threat intelligence report, Xiaomi's MiMo-V2.6 release post and model card, the Artificial Analysis model page and the Hugging Face repository metadata, all read on 22 September 2026.

Now the limits of the claim, which the coverage has mostly skipped. Anthropic's report says the harvested data was used "to strengthen training data used for future models". It does not name MiMo-V2.6, MiMo-V2.5 or any other checkpoint. It could not have named V2.6: the report was published eleven days before the model existed publicly. The report is self-published by a competitor, carries no independent verification, and provides indicators of compromise for other chapters rather than evidence a third party could re-run for this one. That is not a reason to dismiss it. It is a reason to describe it accurately, as a serious first-party allegation about data collection methods, not as a finding about the contents of a specific model.

Anthropic's September 2026 threat intelligence report, case GTG-16008, read against what a procurement file would need

Anthropic statesWhat that establishesWhat it does not establish
More than 400,000 requests across more than 1,500 accounts, over 20 days in March and April 2026A measured volume on Anthropic's own systems, attributed to Xiaomi with high confidenceIndependent confirmation: no third party has audited the attribution
Sessions were replayed to generate data for supervised fine-tuning and reinforcement learningAn alleged collection method and intentThat any released checkpoint contains that data
The data was used to strengthen training data for future modelsAnthropic's characterisation of the purposeWhich model, or whether the data reached one at all
Relayed requests contained names, contact details and corporate data for hundreds of usersA data protection exposure for people who used MiMo through routersWhether any UK data subject was affected, which Anthropic does not address
No indication that Claude's responses were served to Xiaomi's usersA narrower allegation than the ones against Moonshot and DeepSeekThat Xiaomi's own service was otherwise as described to its users

Open source, open weights, MIT: three comforting labels, none of them a control

Xiaomi's release post says "open-sourcing" and calls MiMo-V2.6-Pro "the strongest open-source model to date". Artificial Analysis labels the same model "Open weights". The difference is not pedantry. You can download the parameters. You cannot inspect, audit or reproduce the data that produced them, which is precisely the question the Anthropic allegation raises.

The licence deserves the same scrutiny. The MIT declaration is a metadata tag in the model card's front matter. There is no LICENSE file at the root of the MiMo-V2.6-Pro-RL repository: requests for LICENSE, LICENSE.md and LICENSE.txt all returned 404 on 22 September, and no such file appears in the repository listing. MIT is a permissive licence whose one substantive obligation is to reproduce the copyright notice and permission notice. With no file and no named copyright holder, there is nothing to reproduce, which leaves a buyer relying on a tag rather than on a grant. That tag is more generous than the custom licences on the nearest rivals, GLM-5.3 and Kimi K3, and more generous than the Qwen3.8-Max terms. Generosity is not the issue. Enforceability is.

One more label to check. The 9 billion parameter model Xiaomi offers as a base for further training is a fine-tune of Alibaba's Qwen3.5-9B, which is itself published under Apache 2.0. Xiaomi tags the derivative MIT. Anyone planning to build on that model should read both licences rather than the newer tag, because the terms of the upstream base do not disappear when a downstream card declares something shorter.

Where the tokens go if you do not self-host

Open weights are only a procurement answer if you can afford to serve them. Xiaomi's own SGLang recipe for Pro runs tensor parallel across 16 accelerators over two nodes, with expert parallelism, against roughly 566GB of fp8 weight files. Most organisations will therefore reach for the API, and that is where the governance question becomes concrete.

On 22 September, OpenRouter listed exactly one serving provider for MiMo-V2.6-Pro: Xiaomi itself, at fp8, unmoderated. Using the model through a router does not create a Western-hosted alternative; it adds a second processor in front of the same Chinese endpoint. Anthropic's report makes the same point from the other direction, describing third-party routing services as the path by which hundreds of users' names, contact details and corporate data ended up somewhere they did not expect.

For a UK organisation that means a restricted transfer under the UK GDPR. The ICO's list of jurisdictions covered by UK adequacy regulations runs to the EEA, Andorra, Argentina, the Faroe Islands, Gibraltar, Guernsey, the Isle of Man, Israel, Jersey, New Zealand, South Korea, Switzerland and Uruguay, with partial adequacy for Canada, Japan and the United States under the UK Extension. China is not on it. A transfer to a China-hosted inference API therefore needs an appropriate safeguard such as the International Data Transfer Agreement or the Addendum, plus a completed transfer risk assessment. The specific risk that assessment has to weigh, in this case, is documented in Anthropic's report: an allegation that a lab stored and replayed its own customers' sessions elsewhere.

Three ways to adopt MiMo-V2.6-Pro, assessed against the public record on 22 September 2026

RouteWhat it settlesWhat it leaves open
Self-host the weightsPrompts and outputs stay inside your estate; no restricted transferTraining provenance, licence enforceability, and about 566GB of weights across 16 accelerators
Xiaomi's hosted APILowest cost per task and the configuration the published score was measured onRestricted transfer to China, vendor retention terms, and the Anthropic allegation about stored sessions
A third-party routerConvenience, and one contract covering many modelsAdds a processor without moving the endpoint; Anthropic's report names routers as the exposure path

What to ask before adopting any open-weight model

None of this is a verdict on Xiaomi. The same checklist applies to a Llama derivative, a Mistral release or anything else where the weights are public and the training record is not. The order matters: the cheap questions come first, because most candidates fail them.

Take this with you

Provenance and governance checks, in the order worth doing

  • Find the licence file itself, not the tag. Record the copyright holder, the grant, and whether it covers the weights, the code and the outputs separately.
  • Trace the base model. If the release is a fine-tune, read the upstream licence too and record which terms survive.
  • Ask the vendor, in writing, for a training data provenance statement: pre-training sources, post-training sources, and whether any third-party model outputs were used, with authorisation.
  • Ask specifically whether any teacher model was used in distillation, and which. A card that says multi-teacher without naming the teachers has answered nothing.
  • Record the vendor's answer, or their silence, in the risk register. Silence is a finding, and it dates.
  • Re-run one benchmark that matters to you on the artefact you will actually deploy, at your quantisation and in your harness. Published scores are measured on the vendor's API.
  • For any hosted API, establish the processing location, the retention period, whether prompts are used for training, and the transfer mechanism. Complete a transfer risk assessment where the destination has no UK adequacy regulations.
  • Check whether a router sits in the path, and what it logs. A router changes who holds your prompts, not where they are processed.
  • Decide the data classes allowed through the model before the first pilot, and enforce it at the gateway rather than in a policy document.
  • Set a review trigger: a new allegation, a licence change, or a new checkpoint means the provenance question is reopened, not carried over.

The question that exposes the gap

Xiaomi has published a great deal this week: the weights, a technical report, the reinforcement learning framework, roughly 7,000 graded training tasks, the cost of the run and the exact batch shape. It is a genuinely open engineering disclosure, and it is also a disclosure about the last six days of a training pipeline that ran for far longer. Anthropic has published a specific allegation about part of the period before those six days, and has not connected it to any released model. Both statements can be true at once. Neither is a provenance statement.

So the question to put to any vendor selling you an open-weight model, in writing, is the one neither party has answered here: name every model whose outputs went into this checkpoint, at any stage, and state the authorisation under which they were obtained. A vendor that will not answer is telling you something. A vendor that answers in a leaderboard position is telling you something else.

Sources

  1. PrimaryXiaomi's official MiMo-V2.6 release post, used for the model line-up, the RL run figures, the benchmark claims and the API price listXiaomiaccessed 2026-09-22
  2. PrimaryThe MiMo-V2.6-Pro-RL model card, used for parameter counts, architecture, licence tag, benchmark table, the multi-teacher distillation description and the deployment recipeHugging Faceaccessed 2026-09-22
  3. PrimaryRaw model card front matter, used to confirm that the licence is declared as a metadata tagHugging Faceaccessed 2026-09-22
  4. PrimaryThe 9B distilled model card, used to confirm the Qwen3.5-9B base and the MIT tag on a derivativeHugging Faceaccessed 2026-09-22
  5. PrimaryIndependent ranking page for MiMo-V2.6-Pro, used for the Intelligence Index score, index composition, cost per task, output speed and class rankingArtificial Analysisaccessed 2026-09-22
  6. PrimaryComparison page for Claude Opus 5.5, used for the index score, list price and cost per index taskArtificial Analysisaccessed 2026-09-22
  7. PrimaryAnthropic's September 2026 threat intelligence report, used for case GTG-16008 and the distillation chapter in Anthropic's own wordsAnthropicaccessed 2026-09-22
  8. PrimaryLive endpoint listing for MiMo-V2.6-Pro, used to establish that the only serving provider on 22 September was Xiaomi itself, at fp8OpenRouteraccessed 2026-09-22
  9. PrimaryICO list of countries covered by UK adequacy regulations, used to confirm that China is not among themInformation Commissioner's Officeaccessed 2026-09-22
  10. PrimaryICO guide to international transfers, used for the transfer mechanism and transfer risk assessment requirementsInformation Commissioner's Officeaccessed 2026-09-22
  11. Reported byThe news report that connected the launch to Anthropic's allegation, used as the pointer to both primary sourcesThe Decoderaccessed 2026-09-22
  12. Reported byEarlier coverage of the Anthropic threat report, used to date the report's first public coverageThe Decoderaccessed 2026-09-22
  13. Reported byReport stating that representatives for Xiaomi and the other named labs did not respond to requests for commentBusiness Insideraccessed 2026-09-22

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.