Microsoft calls Decision-1 '4.5 times faster'; JevBench's raw timing for the runner-up gives 1.4 times
Microsoft says Decision-1 is 4.5 times faster than the runner-up and 35 times faster than GPT-6 Sol, and tops a 36-benchmark ranking. The runner-up's 380 ms is a JevBench formula applied to a 116 ms raw timing, and no general-purpose LLM is in the accuracy ranking.
By Parminder Kumar Sharma · · 26 min read

A '4.5 times faster' that turns into 1.4 times, and an accuracy ranking with no general-purpose LLM in it
Microsoft's launch post for Microsoft-Decision-1, published on 9 October 2026, says the model was "the fastest measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol". The post embeds the chart data behind both figures, and they reproduce: 380 ms against 85 ms is 4.47, and 3,010 ms against 85 ms is 35.4 (derived). Those are Microsoft's published numbers. What sits behind each is not the same.
The 85 ms is Microsoft's own median through its Foundry service, which the chart footnote says was "measured through Foundry in the same region". The 380 ms is not a measurement of Quyet. It comes from JevBench, an independent leaderboard, which timed Quyet-1.0-Large at 116 ms (115.8 in its results file) and then applied a formula, twice the raw time plus 150 ms. JevBench's own notes call that adjustment "an assumption, not a hardware normalization". On the raw figure the lead over the runner-up is 1.4 times (derived). Neither figure is like for like: the raw one leaves out a network call that Microsoft's 85 ms includes, and the adjusted one adds a constant that nobody measured.
The accuracy ranking has a second gap. The chart titled "Accuracy, latency, and calibration" ranks Decision-1 first at 83.5 per cent, a mean over 36 benchmarks, ahead of five other decision models from 77.2 to 81.9 and a sixth that answered only 23 of the 36. GPT-6 Sol and Jev 1.13.0 are both listed as "not ranked" for accuracy. So the post's claim to "outperform both LLMs and other decision models" rests, for LLMs, on speed and cost charts and on Microsoft teams saying it was "competitive on quality", with no quality figure printed.
None of this shows that Decision-1 is slow or inaccurate. It shows what kind of evidence the post is. Here is what it does not establish.
- That the speed gap is 4.5 times on your network. Hardware, batch size, concurrency, prompt length and cold starts for the 85 ms are not stated.
- That 35 times is Decision-1 against GPT-6 Sol as you would run it. The reasoning effort, prompt, version and route for Sol are not stated, and OpenAI's API lists two Sol models.
- That the 1.6 point accuracy lead is outside the noise. No per-benchmark scores, question counts or intervals are published.
- That "blind from training" has been verified. It is a statement by the party that built and sold the model.
- That anyone outside Microsoft has repeated it. We found no independent result for Decision-1 on 9 October.
What Microsoft published: the chart data, in one table
Six interactive charts are embedded in the post as encoded documents. We decoded them and read their data arrays as data only. The first, "Accuracy, latency, and calibration", carries everything the headline rests on. The footnotes say accuracy is the "mean score across 36 benchmarks (147,137 questions)", latency is "JevBench v1.6.1 adjusted median (p50), checked Oct 7, 2026" with Decision-1 measured through Foundry, and calibration means 100 is a model whose confidence exactly matches how often it is right.
Microsoft-Decision-1 launch post, 9 October 2026: data array of the chart "Accuracy, latency, and calibration", decoded by us from the page source
| Model | Accuracy, mean of 36 benchmarks | Median latency (p50) | Calibration, 100 is perfect |
|---|---|---|---|
| Microsoft-Decision-1 | 83.5 | 85 ms (p95 125 ms) | 92.2 |
| Quyet-1.0-Large | 81.9 | 380 ms | 93.1 |
| Surogate Rune 26B-A4B | 79.7 | 380 ms | 91.8 |
| GPT-6 Luna Decisions | 79.4 | 300 ms | 89.9 |
| deck-31B | 77.8 | 400 ms | 83.5 |
| H2O-Lightning-4B v1.1 | 77.2 | 210 ms | 91.8 |
| Strands-Decider 2B | 54.8, on 23 of 36 benchmarks | not measured | not scored |
| Jev 1.13.0 | not ranked | 240 ms | not scored |
| GPT-6 Sol | not ranked | 3,010 ms | not scored |
Read the table against the headline. Decision-1 leads on accuracy by 1.6 points over Quyet-1.0-Large and by 3.8, 4.1, 5.7 and 6.3 points over the four models below it (all derived). It is second on calibration, 0.9 behind Quyet, and the post says so in the chart note. It is the fastest row, but the second fastest, H2O-Lightning-4B, is 2.5 times slower on this chart, not 4.5, because the runner-up on accuracy is not the runner-up on speed. The "4.5 times" is the lead over the accuracy runner-up.
The 147,137 questions across 36 benchmarks average 4,087 per benchmark (derived). The post says the benchmarks are "public and private" and gives no list, no per-benchmark scores and no question counts. The mean is described only as a "mean score across 36 benchmarks"; we read it as an unweighted mean of 36 per-benchmark scores, which is an inference, because the post does not say whether it is weighted by question count.
Stated and not stated, claim by claim
What the Microsoft post and the Foundry model card state, and what they do not, read 9 October 2026
| Claim | What Microsoft publishes | What it does not publish |
|---|---|---|
| Highest accuracy | 83.5 against 81.9 for the runner-up, mean of 36 benchmarks, 147,137 questions, among six decision models. | Which 36 benchmarks, which are private, per-benchmark scores, question counts, any interval, who labelled the answers. No general-purpose LLM is in the ranking. |
| Blind from training | Benchmarks "kept blind from training". The card says held-out internal sets "were used to guard against train-test overlap". | The result of any overlap check, which benchmarks are public, what the Qwen3.5-9B base model saw in pretraining, what the rivals trained on. |
| 4.5 times faster | 85 ms p50 through Foundry against 380 ms for Quyet-1.0-Large. | Hardware, concurrency, batch size, tokens and cold starts for the 85 ms. That the 380 ms is a formula applied to a 116 ms raw timing. |
| 35 times faster than GPT-6 Sol | 85 ms against 3,010 ms. A cost chart says both read "the identical brief". | Version, reasoning effort, prompt, route and region for Sol, or where the 3,010 ms was measured. |
| Calibration | Score 92.2, second to Quyet at 93.1. "A 90% prediction should be right about nine times out of 10." | How the score is computed, any reliability curve, calibration by benchmark, language or group. |
| Robustness | Decisions flip on 1.3% of perturbations on average over eight kinds, and on none when options are paraphrased, reversed or shuffled. | The eight kinds, how many items, a baseline for other models. The card says scores "can shift with phrasing and option ordering". |
| Safety | 5,250 requests across 11 benchmarks. The model "successfully refused harmful behavior while retaining a high degree of utility". | Detection rates, false-block rates or any per-benchmark result. |
| Internal tests | Xbox Research sorted over 10,000 items: "competitive on quality" with GPT-6 Sol, over 14 times faster, 200 times cheaper. Copilot: competitive with "GPT5.6 Luna", 100 times faster. | A quality figure in either case. The speed and cost charts cover 2,124 texts in three datasets, not the 10,000. |
| Price | $0.042 per million input tokens, output free. | A price on the model card, which says "The provider has not supplied this information". Quotas, rate limits, an SLA. |
| Availability | In Foundry now, "coming soon through OpenRouter". | A date. OpenRouter's public model list showed no Decision-1 entry at 22:28 BST. |
One row deserves a second look. The post says Decision-1 "achieved the highest accuracy". The Foundry model card, Microsoft's own document for customers, says "Overall, it performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology". Both can be true of the same numbers. They are not the same claim, and a buyer should quote the second.
The 4.5 times: a measured API against a formula
JevBench is a public leaderboard for decision models, run by Benchmark Heaven, which says it is maintained independently of TypeSafe AI, the company behind the Jev model. Its own pages call it a hobby project run by a one-person company, which also offers custom evaluations and consulting. The v1.6.1 release answers 1,500 decisions per system, 1,200 of them sealed, and ranks 135 systems from a roster of 173.
For open-weights models JevBench ran the model itself on hardware it rented. It publishes two latency figures for each: a raw median and an "adjusted" median, defined as twice the raw time plus 0.15 seconds. The results file labels this "x2 + 0.15 s (assumption, not measured)", and a note says the adjustment "is an assumption, not a hardware normalization". We checked the arithmetic on four rows and it holds exactly: Quyet-1.0-Large 115.8 ms raw becomes 381.6 ms adjusted. Microsoft's chart shows 380. Hosted APIs such as Jev 1.13.0 are not adjusted. JevBench also says a hosted endpoint "chooses its own hardware and sees the benchmark inputs", so it ranks hosted APIs on a separate board. Microsoft's chart puts both kinds in one row set, and sets a hosted API Microsoft measured itself against self-hosted estimates.
Read the chart from the dashed line. On JevBench's raw figures three of the four self-hosted models are 1.4 to 1.5 times slower than 85 ms, and H2O-Lightning-4B is 2.9 times faster, at 29 ms. H2O.ai's own model card reports a 34 ms median on a single GPU, which is the vendor's figure on different hardware. On the adjusted figures every self-hosted model is at least 2.4 times slower. The truth for a given deployment lies somewhere else again, because the raw timings exclude the network call that Foundry's 85 ms includes, and the adjustment is a flat guess. A reader who wants "fastest" can say 85 ms is the lowest figure on the chart. A reader who quotes "4.5 times" is quoting a ratio of a measured figure to a modelled one.
The formula also explains why the accuracy runner-up looks so slow. Of Quyet's 381.6 ms, 150 ms (39 per cent) is the added constant (derived).
What does the gap mean in time? The post's own example is that adding 100 ms to each of 20 sequential decisions adds two seconds to a workflow. Using the chart's medians, a chain of dependent decisions waits like this.
Median latency from Microsoft's chart multiplied by the number of dependent decisions (derived; ignores queueing, retries and the work between decisions)
| Model | One decision | 20 in sequence | 100 in sequence |
|---|---|---|---|
| Microsoft-Decision-1 | 0.085 s | 1.7 s | 8.5 s |
| H2O-Lightning-4B v1.1 | 0.21 s | 4.2 s | 21 s |
| GPT-6 Luna Decisions | 0.30 s | 6.0 s | 30 s |
| Quyet-1.0-Large | 0.38 s | 7.6 s | 38 s |
| GPT-6 Sol | 3.01 s | 60 s | 301 s |
So the speed gap matters most in two places: an interactive path where a person waits (a routing step under a second against one near three seconds), and a long agent chain of dependent calls. For a nightly job that labels survey comments, 85 ms against 380 ms is invisible, and cost and accuracy decide. Microsoft's own Xbox example makes the same point: its speed chart shows 143 to 188 ms per text for Decision-1 against 2.6 to 2.8 seconds for Sol.
The 35 times, and which GPT-6 Sol
Microsoft writes "GPT-6 Sol" with no version, snapshot or setting. OpenAI's API lists gpt-6-sol and gpt-6.1-sol as separate models, both at $2.00 input and $10.00 output per million tokens. The gpt-6-sol page says "See GPT-6.1 Sol for the newer Sol model", supports reasoning effort from none to max, and defaults to medium. We read Microsoft's "GPT-6 Sol" as the model ID gpt-6-sol, which is an inference from the name. The site's earlier briefs show why the name is not enough: GPT-6 Sol is already two versions under one name in ChatGPT and Codex, and GPT-6.1 Sol is a different model again.
The effort matters more than the version, because Sol thinks before it answers and the chart does not say how long. Artificial Analysis, an independent measurer, publishes Sol's median time to first token by setting. Its figures use its own test prompt through OpenAI's first-party API, not Microsoft's brief, so they show the range rather than reproduce the 3,010 ms.
Artificial Analysis, GPT-6 Sol median time to first chunk by reasoning effort, read from its model pages on 9 October 2026; the multiple of 85 ms is ours
| Reasoning effort | Time to first chunk | Multiple of 85 ms |
|---|---|---|
| None | 0.94 s | 11x |
| Low | 1.72 s | 20x |
| Medium (API default) | 2.78 s | 33x |
| High | 26.6 s | 312x |
| Max | 154.7 s | 1,820x |
Microsoft's 3,010 ms sits just above Artificial Analysis's medium figure and far below its high one, which fits a short prompt at about the default setting (inference). If a team runs Sol with reasoning off for a classification task, the speed gap shrinks towards 11 times on that independent prompt. The post does not say what task framing Sol was given: the prompt, the effort, whether a structured output was requested. A cost chart notes that both models "read the identical brief, so input tokens match; only GPT-6 Sol also pays for output". If both read the same 262 or so input tokens (derived from the $11 per million texts), roughly 78 per cent of Sol's bill in that chart is output tokens (derived, on those assumptions), which is where a reasoning setting would show up (inference).
It also is not like for like in kind. Sol generates; Decision-1 reads once and returns probabilities for options you supply. The Foundry card says the model "does not generate explanations or rationales" and is "not designed for text generation, open-ended question answering, conversation, translation, or summarization". Beating an LLM at closed questions says nothing about reasoning or writing, and Microsoft does not claim it does. Independent evidence on this trade-off exists, from a different benchmark: on JevBench's hosted-API board, GPT-6 Luna as a general model scores about 97 on chance-corrected intelligence against about 57 for OpenAI's Decisions API, which takes gpt-6-luna as its model, at 1.82 s against 0.30 s and a modelled $0.113 against $0.052 per 1,000 decisions. An analytics team that tested a decision model against LLMs found the LLMs more accurate in its own synthetic test. Neither says anything about Decision-1. Both say speed and price are bought with accuracy at the category level, and that the accuracy side is the one Microsoft's ranking leaves general-purpose LLMs out of.
List prices read on 9 October 2026, and what a million decisions costs at the token count implied by Microsoft's cost chart (derived)
| Route | Input per million tokens | Output per million tokens | Per million decisions |
|---|---|---|---|
| Microsoft-Decision-1 (Microsoft post) | $0.042 | free | about $11 |
| OpenAI Decisions API, gpt-6-luna (OpenAI guide) | $0.10 | not charged | about $26, same tokens assumed |
| GPT-6 Sol (OpenAI model page) | $2.00 | $10.00 | about $2,434, Microsoft chart estimate |
At 100,000 decisions a month the saving over Sol is about $240 and over OpenAI's Decisions API about $1.50 (derived, USD list prices, same-token assumption for the second). At 10 million a month it is about $24,000 and $150. Price per call is real, and for most UK teams staff time, evaluation and review capacity will cost more than the model.
The accuracy lead: 1.6 points on a mean of 36
The headline accuracy figure is a mean of per-benchmark scores. Three checks apply. The first is our calculation from the numbers above, not Microsoft's.
Is 1.6 points more than noise? Pooled over 147,137 questions, a binomial 95 per cent interval around 82 per cent is about 0.2 points wide on each side. But the questions are not independent draws, and the score is a mean of 36 benchmark means. A single benchmark of 200 questions has an interval of about 5.3 points either side at 82 per cent, 500 questions 3.4, 1,000 questions 2.4. For a mean of 36 paired differences, the lead clears 1.96 standard errors only if the benchmark-to-benchmark spread of the difference is below about 4.9 points (derived). The post publishes no per-benchmark scores, so that cannot be checked. It is plausible and not shown.
Who chose the field? Microsoft says it took "several of the top public models" on JevBench. In the board version we read on 9 October, JevBench's open-weights Capability list also had decisio v0.8.0 at second, René-1 31B FP8 at fourth and Bobcat Flash 1.2 at sixth, none of which is in the chart, while Jev 1.13.0, the reference system JevBench says "defines the genre", appears for latency and is "not ranked" for accuracy. Strands-Decider 2B is averaged over 23 of 36 benchmarks, 13 fewer, so its 54.8 is not comparable. The "#n on JevBench" labels also do not all come from one list: three match the open-weights order, one matches the hosted-API order, and one is off by a place against the version read.
What does "blind" mean? The training disclosure says the post-training data are "publicly available datasets" processed under Microsoft's open-data process, plus synthetic data, "first used in September 2026", with "data collection is ongoing" and no training cut-off date supplied. Public human-labelled datasets were "prioritized". A benchmark that is public can be in someone's training data; the card says held-out internal sets were used to guard against overlap, and does not report what was found. The post cannot say what the Qwen3.5-9B base model or the rivals saw. A public community suite exists at similar scale (the Surogate Rune card describes a Decision Index suite of 151,034 requests from 44 benchmarks); Microsoft does not say whether its 36 overlap with it. That is a question for Microsoft, not a finding.
Calibration is the property the model is sold on, and the post's evidence is one number per model. Quyet scores higher (93.1 against 92.2), and a score of 92 does not by itself say how far a 0.9 probability can be trusted on your tickets. The card itself says "Calibration is strongest on familiar task types" and lists non-English languages, specialised domains (medical, legal, financial) and ambiguous options as weak areas. The model is optimised for English and was evaluated in 22 other languages named on the card.
Who the rivals are, and who has what interest
Most readers will not have heard of the runner-up. The six models in Microsoft's appendix are mostly small open-weights projects, and two of them are a single developer's work.
The models in Microsoft's appendix, from each model's own page, read 9 October 2026
| Model | Maker and licence | What its page says |
|---|---|---|
| Quyet-1.0-Large | An individual developer. Apache-2.0. | Gemma-4-31B-it with a LoRA fine-tune, 31.3 billion parameters, 62.5 GB weights, one 80 GB GPU. 503 downloads last month. Returns calibrated probabilities. |
| Surogate Rune 26B-A4B | Invergent. Apache-2.0, gated download. | 26.5 billion parameters, 8 of 128 experts active per token. Card claims second place on a Decision Index 0.2.1 board, 0.45 behind Jev 1.13. |
| GPT-6 Luna Decisions | OpenAI. Hosted API, public beta. | The Decisions API, "10x faster than the Responses API", gpt-6-luna only, $0.10 per million input tokens, Zero Data Retention for eligible customers, residency in the US and Europe (EEA plus Switzerland). |
| deck-31B | An individual developer. Apache-2.0 server, Gemma terms for the weights. | Unmodified Gemma-4-31B-it with a prompt recipe: "We did not train them". Four stars on GitHub. |
| H2O-Lightning-4B v1.1 | H2O.ai. Apache-2.0. | Qwen3.5-4B base. Card claims first place on JevBench Composite on 7 October and 34 ms median on one GPU. Says no JevBench data was used in training. |
| Strands-Decider 2B | strands-labs, labelled AWS in the chart. Apache-2.0. | Qwen3.5-2B base, 541 GitHub stars. Answered 23 of the 36 benchmarks. |
Each party in this story has an interest, and the rule is to say so plainly. Microsoft sells Decision-1 through Foundry, chose the comparison set, ran the Foundry timing and wrote every headline figure in the post. OpenAI sells both GPT-6 Sol and the GPT-6 Luna Decisions API that the chart ranks fourth on accuracy at 79.4, so it has its own stake in how a rival decision model compares; Microsoft, for its part, held an investment in OpenAI Group PBC that Microsoft's own post of 28 October 2025 put at roughly 27 per cent (as of that post; not re-checked), and says it will rebase Decision-1 on OpenAI models too. Quyet-1.0-Large is an open-weights model whose card names no price or commercial service, so no one is speaking for the runner-up in this comparison, which also means no one can correct it. JevBench's operator offers custom evaluations and consulting, an interest of a different and smaller kind.
The category is not new, whatever the label suggests. OpenAI's Decisions API is in public beta and JevBench ranks 135 systems. Microsoft's contribution is a Microsoft-hosted, enterprise-sold entry, which may matter to buyers who can only procure through suppliers they already use (inference).
What the model is, and the terms it is sold on
The label is doing work. "Decision model" and "decision intelligence" suggest something that decides. The Foundry card says what it is: Alibaba's open-weight Qwen3.5-9B, a dense decoder-only transformer, "post-trained to return calibrated probability scores over fixed answer options". Given a situation, a question and a fixed set of options, it returns a probability for each. It does not explain. The card puts the decision where it belongs: "The integrating application is responsible for defining appropriate answer options, confidence thresholds, escalation paths, human oversight, and safeguards for downstream actions." So the model scores, and whoever sets the options and the threshold decides. That is a useful classifier with a good interface, not a decision-maker, and "workflow control" is a sentence about what you may wire it to, not a claim that it controls anything safely.
The post's split between "LLMs" and "decision models" is also looser than it sounds. Every row in the chart, Decision-1 included, is a language model wired to read once and output probabilities (Qwen3.5-9B, Gemma-4-31B, gpt-6-luna and others). The split is about how a model is used, not what it is.
Microsoft Foundry model catalogue entry and model card for Microsoft-Decision-1, read 9 October 2026 in a browser and from the page source
| Item | What the card states | Not stated or qualified |
|---|---|---|
| Input and output | Text only, up to 32,768 input tokens, a fixed set of options. Output is JSON probabilities. | A sample JSON response ("The provider has not supplied this information"). |
| Hosting | Decision-1 "is available through a hosted API in Microsoft Foundry". Model weights are not distributed. | Any way to run it in your own tenancy or offline. |
| Status and terms | Lifecycle: "Generally available (GA)". | The licence text points to the "preview supplemental terms" for Azure, so the label says GA and the terms link says preview. |
| Price | None on the card. The post says $0.042 per million input tokens. | The "View pricing" link led to an Azure pricing page that does not mention the model. |
| Regions | The catalogue data lists one deployment type, Global Standard, with the location code WESTCENTRALUS. | Any Data Zone, EU or UK deployment. Learn pages on regions and quotas do not mention the model. |
| Training data | Public datasets under Microsoft's open-data process plus synthetic data; "No Microsoft customer data was used". Training time September to October 2026. | The training cut-off date and the dataset list. The size is given only as a range, 1 billion to 1 trillion tokens. |
| Languages | Optimised for English, evaluated in 22 other languages. | Per-language accuracy or calibration. "Coverage, quality, and calibration may vary". |
| Evolution | Version 1. Microsoft will "rebase" on MAI and OpenAI models and keep releasing updates. | How a customer is told, whether a version is pinned, how long it is kept. |
| Dates | Card: model release 8 October 2026. Post: 9 October. | Why they differ; probably the announcement lagging the release (inference). |
Nothing independent yet
We looked for a result for Decision-1 that Microsoft did not produce. As of about 22:30 BST on 9 October 2026 we found none.
- JevBench, the board Microsoft cites, has no Decision-1 row on its v1.6.1 results file or its v1.6.2 page as fetched.
- Artificial Analysis and Vals AI listed no Decision-1 entry on the pages we read. Both focus on general models, so the absence says little.
- OpenRouter's public model list (458 models, read at 22:28 BST) had no Decision-1 entry and no decision model from Microsoft, so "coming soon" is unverified.
- Press coverage repeats the post. Four trade pieces we read say the figures are Microsoft's own, and one says the announcement does "not provide the full datasets or latency test conditions needed to reproduce the comparison independently". Two further articles refused our request (HTTP 403) and were not read or worked around.
- Microsoft's Learn pages on models sold by Azure, regions, deployment types and quotas do not mention Decision-1.
What does exist is independent evidence about the category, not about this model. An Amplitude engineering analysis of Jev says decision models share "the same pitfalls in accuracy" as LLMs, calibrate differently on different data, and fail confidently on near-synonym intents: the type system constrains the shape of the answer, not the judgment. JevBench's methodology says "confidence, coherence, robustness and real-world error costs need additional workload-specific checks". Both are reasons to run your own test before relying on anyone's 36.
UK: what to put on paper before anyone connects it
Model change control. The catalogue entry reads "Version: 1". Microsoft says it will rebase Decision-1 on other models and keep releasing updates, and nothing we read says how customers are told or whether an earlier version stays available. Treat a rebase as a new model with a new evaluation and a new inventory entry. The site's earlier brief on one GPT-6 Sol name covering two versions is the same lesson from a different vendor, and the price-cut brief shows what a vendor-run comparison usually leaves out.
Where the data goes. The catalogue's page data lists one deployment type, Global Standard. Microsoft's deployment-types page (last updated 12 August 2026) says Global types "May be processed in any Azure region", while Data Zone types stay in the US, EU or Asia Pacific zone. No Data Zone or UK region was found for Decision-1. Microsoft's data-privacy page for models sold by Azure (last updated 19 May 2026) says prompts and completions are not used to train or improve the base models and the models are stateless, with abuse monitoring applied. Whether that page applies unchanged to a model released on 8 October is not stated. For comparison, OpenAI's Decisions guide names the United States and Europe (EEA plus Switzerland) for residency and does not name the UK. If prompts carry personal data, ask your data protection officer about the location of processing and the contract before testing with real data; the card names Microsoft Ireland Operations Limited as "Authorized representative", and a UK representative is not stated.
Decisions about people. If a score decides who is called back, whose complaint is escalated, which candidate is read or which employee is flagged, Articles 22A to 22D of the UK GDPR may apply. They were substituted by the Data (Use and Access) Act 2025 and, per legislation.gov.uk, took effect from 5 February 2026 so far as not already in force. Article 22A treats a decision as based solely on automated processing if there is "no meaningful human involvement", and as significant if it has a legal or similarly significant effect on the person. Article 22C then requires safeguards: information about the decision, a way to make representations, a way to obtain human intervention, and a way to contest it. The ICO's guidance, a draft updated on 31 March 2026 and under consultation, says a human should assess the decision at an appropriate point, be able to influence and alter the outcome, be trained to understand the system's outputs and limits, and take the relevant data into account, on non-exhaustive criteria. It says ad hoc spot-checking is not enough and a record should be kept. Microsoft's card says Decision-1 is "not designed or evaluated for use as the sole automated decision-maker in consequential decisions about people" and names credit, employment, housing, insurance, education, healthcare and legal rights. The site's earlier briefs on AI call scoring at a UK law firm and on the ICO's work with AI developers cover how scores become decisions.
What a score replaces. The card suggests you "automate high-confidence outcomes and escalate low-confidence or ambiguous outcomes for additional review". That is a design with three parts: a threshold, a queue and a person. Before connecting it, write down what the step does today (a human triager, a rule, an LLM), who owns the threshold, who works the escalation queue and how many items a day that queue can take. A score that quietly empties the queue has replaced a person, and the replacement is the part that needs evidence.
Fairness and calibration. The card says scores "may reflect biases from base model and training data" and should not be the sole basis for decisions about individuals. It lists counterfactual bias tests with varied demographic details among the evaluations and prints no result. The post says nothing on fairness. Calibration is one number averaged over 36 benchmarks, and the card says it is strongest on familiar task types.
Security. A model that screens requests or proposed agent actions reads untrusted text and sits in a control path. The post says it tested jailbreak and prompt-injection requests, and prints no detection or false-block rate. The NCSC's guidance on adopting agentic AI says humans remain accountable for the decision to deploy, the access granted, the safeguards and the consequences, and that an agent should never get unrestricted access to sensitive data or critical systems. Its secure-design guideline asks that choosing an external model provider includes due diligence on that provider's security posture. A score that gates an agent is one safeguard to evidence, not a substitute for least privilege.
What to do, in order
Take this with you
Before a decision model goes near production
- Inventory where one would sit: every place a person, a rule or an LLM routes, classifies, ranks, verifies or approves today, and mark the ones that touch personal data.
- Decide what it may never decide alone: anything with a legal or similarly significant effect on a person, and any action that cannot be undone.
- Build your own labelled test set of 500 to 1,000 items from your own data, with a written rubric and two labellers, and keep a sealed part you never tune on. At 82 per cent accuracy, 500 items give an interval of about 3.4 points either side and 1,000 about 2.4.
- Measure accuracy and error type per class, not the average, and price each error type. Work out the cost per 1,000 decisions from your own token counts.
- Test latency on your own network, region and concurrency at p50 and p95, with cold starts, against the model you use now at the same settings.
- Check calibration before using scores as thresholds: bin your results, compare stated confidence with the hit rate, and set thresholds from the cost of errors. Include a "cannot tell" option.
- Log inputs, outputs, scores, thresholds and the model version, and keep them for the period your data protection officer sets.
- Pin the version and treat any rebase as a new model. Record it in the model inventory with an owner, purpose, data, region and evaluation date, including in an ISO/IEC 42001 management system if you run one.
- Plan and test a fallback for timeouts and outages: another model, a rule set or a person.
- Set a re-evaluation date, 90 days is a sensible start, and re-run on any model change. Ask Microsoft for per-benchmark results, the list of benchmarks and the latency test conditions.
The question the post cannot answer
If a score from a model you did not train, tested on benchmarks you cannot see, decides which cases a person never reads, who has measured its errors on your cases, with what labels, and when did they last look?
Key facts
Sources
- PrimaryIntroducing Microsoft-Decision-1, our model for fast decision-making, 9 October 2026: the whole post read from the HTML and a second fetch at 22:17 BST, plus the six embedded chart documents decoded and read as data (accuracy, latency and calibration; latency; average accuracy; calibration; speed per classification; cost per run)Microsoft Command Lineaccessed 2026-10-09
- PrimaryModel catalogue entry and model card for Microsoft-Decision-1 read in a browser tab and from the page source: capabilities, out-of-scope uses, technical specs, training disclosure, evaluation and benchmarking methodology, responsible AI, known limitations, licence text, quick factsMicrosoft Foundryaccessed 2026-10-09
- PrimaryJevBench v1.6.1 open-weights leaderboard, read for the Jev-class capability ranking, the hosted-versus-self-hosted fairness note and the operator's own descriptionBenchmark Heavenaccessed 2026-10-09
- PrimaryJevBench v1.6.1 published aggregate results file (JSON): raw and adjusted p50 latency per system, the x2 + 0.15 s adjustment label and note, sample sizes, ranked and roster counts; used to check Microsoft's 380 ms for Quyet-1.0-LargeBenchmark Heavenaccessed 2026-10-09
- PrimaryJevBench hosted API leaderboard: OpenAI Decisions and GPT-6 Luna rows (intelligence, calibration, latency, modelled cost); no Microsoft-Decision-1 or GPT-6 Sol rowBenchmark Heavenaccessed 2026-10-09
- PrimaryJevBench methodology, limitations and operator description, read for the capability formula, the cost and latency caveats and the one-person-company footerBenchmark Heavenaccessed 2026-10-09
- PrimaryDecisions API guide (public beta): question types, gpt-6-luna as the only model, input-only price of $0.10 per million tokens, Zero Data Retention, data residency in the US and EuropeOpenAIaccessed 2026-10-09
- PrimaryGPT-6 Sol model page: model ID, reasoning effort options and default, price of $2 input and $10 output per million tokens, pointer to GPT-6.1 Sol as the newer Sol modelOpenAIaccessed 2026-10-09
- PrimaryAPI pricing page: gpt-6-sol, gpt-6.1-sol, gpt-6-luna and gpt-5.6-luna rowsOpenAIaccessed 2026-10-09
- PrimaryGPT-6 Sol (Medium) page: median time to first chunk 2.78 s on its own test prompt; the other effort variants (none, low, high, max) read from their own pages for the tableArtificial Analysisaccessed 2026-10-09
- PrimaryQuyet-1.0-Large model card: base model, size, licence, hardware, download count; the runner-up's own descriptionHugging Faceaccessed 2026-10-09
- PrimarySurogate Rune 26B-A4B v3 model card: licence, size, the Decision Index 0.2.1 table and its 151,034-request, 44-benchmark suite descriptionHugging Faceaccessed 2026-10-09
- PrimaryH2O-Lightning-4B model card: Qwen3.5-4B base, licence, first-place JevBench claim of 7 October, 34 ms median latency on one GPU, training-data disclosureHugging Faceaccessed 2026-10-09
- Primarystrands-decider repository README: base model, licence, star count, example outputGitHubaccessed 2026-10-09
- PrimaryDeployment types for Foundry Models, last updated 12 August 2026: where Global and Data Zone deployments process data; no mention of Decision-1Microsoft Learnaccessed 2026-10-09
- PrimaryData, privacy and security for Foundry Models sold by Azure, last updated 19 May 2026: stateless models, no training on prompts, abuse monitoring, location of processingMicrosoft Learnaccessed 2026-10-09
- PrimaryThe next chapter of the Microsoft and OpenAI partnership, 28 October 2025: Microsoft's investment in OpenAI Group PBC of roughly 27 per cent as of that postMicrosoftaccessed 2026-10-09
- PrimaryOpenRouter public model list, 458 models at 22:28 BST on 9 October 2026: no Microsoft-Decision-1 entryOpenRouteraccessed 2026-10-09
- PrimaryUK GDPR Article 22A, solely automated processing and significant decisions, with the amendment note on commencement; Articles 22B, 22C and 22D read alongsidelegislation.gov.ukaccessed 2026-10-09
- PrimaryUK GDPR Article 22C, safeguards for automated decision-making: information, representations, human intervention, contestinglegislation.gov.ukaccessed 2026-10-09
- PrimaryICO draft ADM guidance updated 31 March 2026: when the provisions apply, what meaningful human involvement means, and that ad hoc spot-checking is not enoughInformation Commissioner's Officeaccessed 2026-10-09
- PrimaryICO draft ADM guidance, safeguards page: information about decisions, representations, human intervention and contestingInformation Commissioner's Officeaccessed 2026-10-09
- PrimaryThinking carefully before adopting agentic AI, 15 May 2026: human accountability, least privilege, never granting unrestricted accessNational Cyber Security Centreaccessed 2026-10-09
- PrimaryGuidelines for secure AI system development, secure design: due diligence on external model providersNational Cyber Security Centreaccessed 2026-10-09
- Reported byAn analysis of Jev: faster and cheaper, but watch the accuracy: independent engineering analysis of decision models in general, not of Decision-1Amplitudeaccessed 2026-10-09
- Reported byTrade coverage of the launch, read to check for independent results; it says the figures are Microsoft's own measurementsFourWeekMBAaccessed 2026-10-09


