Haiku 5.5 'costs around 75% less', but at Max effort three of Anthropic's four charts show it dearer than 4.5
Anthropic's launch page says Haiku 5.5 costs 90 per cent less per token and around 75 per cent less to run, and footnote 2 explains the gap. Its own chart data show the 75 per cent holds at the default Medium effort and reverses at Max, the setting behind the benchmark table.
By Parminder Kumar Sharma · · 29 min read

Token prices fall 90 per cent, the bill 'around 75 per cent', and footnote 2 holds the gap
Anthropic's launch page for Claude Haiku 5.5, published on 7 October 2026, puts two cost claims side by side. The price per token is 90 per cent lower than Haiku 4.5's for prompts up to 100,000 tokens: input falls from $1.00 to $0.10 per million tokens and output from $5.00 to $0.50. And the model "now costs around 75% less to run". Footnote 2 explains the difference. The 90 per cent applies up to 100,000 tokens and the cut is 50 per cent above that. On Haiku 4.5, the footnote says, 90 per cent of requests fell in the lower band. And the calculation "also accounts for" an updated tokenizer that "uses slightly more tokens per task".
Our arithmetic (derived) reproduces most of the gap but not all of it. Weighting by request count, 0.9 times 0.10 plus 0.1 times 0.50 is 0.14 of Haiku 4.5's price, 86 per cent lower. Anthropic's migration guide puts the tokenizer change at about 30 per cent more tokens for the same text, which lifts 0.14 to 0.18, 82 per cent lower. The page says 0.25. It prints neither the weights, the token factor nor the workload, so the last step cannot be reproduced from the page. One combination that closes it is 23 per cent of Haiku 4.5's spend, rather than its requests, sitting on prompts over 100,000 tokens. That is our inference, and a plausible one, because a long prompt costs far more than a short one. A reader who divides 25 by 10 and concludes "2.5 times the tokens per task" has misread the footnote: the 90 per cent never applied to every request.
The checkable fact is in the charts. The launch page embeds the numbers behind its accuracy against cost charts, and the benchmark table's scores are at Max effort. At Max, Haiku 5.5 costs more per task than Haiku 4.5's plotted point on three of the four charted benchmarks: 3.6 times on GDPval-AA, 3.3 times on Terminal-Bench 4.0 and 2.3 times on Humanity's Last Exam, against 0.42 times on OSWorld (all derived). At the API default, Medium, all four are cheaper, by 91, 87, 93 and 14 per cent. The "around 75%" sits near the default and a long way from the setting that produced the scores in the table.
That does not establish that the 75 per cent is wrong. It establishes what it is: a vendor average, over a mix of prompt lengths and settings the page does not print. Here is what else the page does not establish.
- That your bill falls by 75 per cent. Your mix of prompt lengths, effort setting and thinking is not Anthropic's. Haiku 4.5 had no effort setting. Haiku 5.5 thinks by default.
- That cheaper per token means cheaper per task. Per token, Haiku 5.5 is a twentieth of Sonnet 5.5's price up to 100,000 tokens. Per attempt on Terminal-Bench, Anthropic's own chart shows a Sonnet point that costs no more and scores higher for every Haiku point except Low.
- That "most capable small model" is a market claim. The launch page says "ever released" and the news page says "yet". The comparison is with Anthropic's earlier models, plus one rival, GPT-6 Luna.
- That a score against zero is a rate. Haiku 4.5's 0.0 per cent on Terminal-Bench 4.0 means it passed none of 660 trials. It defines no ratio.
- That the scores are independent. Most were run by Anthropic. Where outside parties ran them, their figures agree on direction and differ on level.
What Anthropic released, and what the pages leave out
The facts below come from the launch page, the system card, the models overview, the migration guide, the prompting guide and the pricing and effort documentation, all read on 8 October 2026.
Claude Haiku 5.5 at release. Sources: Anthropic launch page, models overview, migration guide, pricing and effort documentation, 8 October 2026
- Item
- Release and ID
- Published
- 7 October 2026. claude-haiku-5-5, a fixed ID with no date suffix and no alias. Retirement not sooner than 7 October 2027
- Not stated or qualified
- There is no dated snapshot to pin. Pinning means pinning the ID and logging what answered
- Item
- Limits and price tiers
- Published
- 1M-token context, 128K output, knowledge cutoff June 2026. $0.10 input and $0.50 output per million tokens up to 100,000-token prompts, $0.50 and $2.50 above
- Not stated or qualified
- Every other Claude model from 4.6 up carries the 1M window at one price. Haiku 5.5 is priced by prompt length (pricing documentation)
- Item
- Effort
- Published
- The first Haiku with effort: Low, Medium, High, Xhigh, Max. Medium is the default on the API and in Claude Code
- Not stated or qualified
- The launch table is at Max. The docs say to use Xhigh or Max only where your own evals show a gain, and to compare Sonnet 5.5
- Item
- Platforms
- Published
- Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS (docs). The launch page says AWS, Google Cloud and Microsoft Azure
- Not stated or qualified
- Bedrock lists Haiku 5.5 as See Access, not among the models open to every customer. Check the console
| Item | Published | Not stated or qualified |
|---|---|---|
| Release and ID | 7 October 2026. claude-haiku-5-5, a fixed ID with no date suffix and no alias. Retirement not sooner than 7 October 2027 | There is no dated snapshot to pin. Pinning means pinning the ID and logging what answered |
| Limits and price tiers | 1M-token context, 128K output, knowledge cutoff June 2026. $0.10 input and $0.50 output per million tokens up to 100,000-token prompts, $0.50 and $2.50 above | Every other Claude model from 4.6 up carries the 1M window at one price. Haiku 5.5 is priced by prompt length (pricing documentation) |
| Effort | The first Haiku with effort: Low, Medium, High, Xhigh, Max. Medium is the default on the API and in Claude Code | The launch table is at Max. The docs say to use Xhigh or Max only where your own evals show a gain, and to compare Sonnet 5.5 |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS (docs). The launch page says AWS, Google Cloud and Microsoft Azure | Bedrock lists Haiku 5.5 as See Access, not among the models open to every customer. Check the console |
The price table, and what each ratio is
Anthropic's price table sets Haiku 5.5 beside Haiku 4.5 and Sonnet 5.5. The Sonnet column shows the new cache-read price: Anthropic halved it from $0.20 to $0.10 on 7 October. Our 28 September Sonnet briefing printed $0.20, which was right that day.
Price per million tokens as printed on the launch page, 7 October 2026. Cache writes are the five-minute rate
- Per million tokens
- Input
- Haiku 5.5: up to / over 100k prompts
- $0.10 / $0.50
- Haiku 4.5 / Sonnet 5.5
- $1.00 / $2.00
- Per million tokens
- Output
- Haiku 5.5: up to / over 100k prompts
- $0.50 / $2.50
- Haiku 4.5 / Sonnet 5.5
- $5.00 / $10.00
- Per million tokens
- Cache reads
- Haiku 5.5: up to / over 100k prompts
- $0.01 / $0.05
- Haiku 4.5 / Sonnet 5.5
- $0.10 / $0.10
- Per million tokens
- Cache writes
- Haiku 5.5: up to / over 100k prompts
- $0.125 / $0.625
- Haiku 4.5 / Sonnet 5.5
- $1.25 / $2.50
| Per million tokens | Haiku 5.5: up to / over 100k prompts | Haiku 4.5 / Sonnet 5.5 |
|---|---|---|
| Input | $0.10 / $0.50 | $1.00 / $2.00 |
| Output | $0.50 / $2.50 | $5.00 / $10.00 |
| Cache reads | $0.01 / $0.05 | $0.10 / $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 / $2.50 |
The ratios (derived). Against Haiku 4.5, every line is 90 per cent lower up to 100,000 tokens and 50 per cent lower above it. Against Sonnet 5.5, input, output and cache writes are 5 per cent of the price up to 100,000 tokens and 25 per cent above, and cache reads are 10 per cent and 50 per cent. OpenAI's page for GPT-6 Luna, the table's one rival, lists $0.10 input and $0.50 output, the same as Haiku 5.5's lower band. The page read prints a single price and states no long-prompt tier.
Anthropic says halving Sonnet's cache reads lowers its cost on most agentic work by around 20 per cent. On the Terminal-Bench chart data the cut is 21 to 26 per cent (derived). One oddity: the chart's "before" line puts Sonnet 5.5 at Max at $14.15 an attempt, while our 28 September briefing read $12.54 from Sonnet's own launch page. Anthropic does not say why the two differ.
The 75 per cent, taken apart
How far each step takes Haiku 4.5's price. Our arithmetic from the launch page and the migration guide
- Step
- Price, prompts up to 100,000 tokens
- Result
- 0.10 of Haiku 4.5
- Basis
- Price table
- Step
- Price, prompts over 100,000 tokens
- Result
- 0.50 of Haiku 4.5
- Basis
- Price table
- Step
- Blend by request count: 90% short, 10% long
- Result
- 0.14 (86% lower)
- Basis
- Footnote 2 weights, derived
- Step
- Plus about 30% more tokens for the same text
- Result
- 0.18 (82% lower)
- Basis
- Migration guide, derived
- Step
- Anthropic's figure, 'around 75% less'
- Result
- 0.25 (75% lower)
- Basis
- Basis not printed
- Step
- Share of spend over 100k that closes the gap
- Result
- 23% of spend
- Basis
- Our inference
| Step | Result | Basis |
|---|---|---|
| Price, prompts up to 100,000 tokens | 0.10 of Haiku 4.5 | Price table |
| Price, prompts over 100,000 tokens | 0.50 of Haiku 4.5 | Price table |
| Blend by request count: 90% short, 10% long | 0.14 (86% lower) | Footnote 2 weights, derived |
| Plus about 30% more tokens for the same text | 0.18 (82% lower) | Migration guide, derived |
| Anthropic's figure, 'around 75% less' | 0.25 (75% lower) | Basis not printed |
| Share of spend over 100k that closes the gap | 23% of spend | Our inference |
Two things sit outside that arithmetic. The first is thinking. Haiku 4.5 ran without it unless you set a budget. Haiku 5.5 thinks by default, and the prompting guide says that telling it to answer directly "didn't stop it from thinking". Thinking can be switched off at Low, Medium and High but not at Xhigh or Max. In Anthropic's own testing, moving from Low to Medium roughly halved early stopping and "more than doubled the output tokens for each attempt".
The second is verbosity. Artificial Analysis counted 440 million output tokens for Haiku 5.5 at Max across its Intelligence Index, against a median of 100 million for the models it groups it with, 4.4 times (derived). Neither effect is in a per-token price.
On Anthropic's own chart data, the 75 per cent lives at Medium
The launch page plots cost per attempt against score at each of Haiku 5.5's five effort levels for three benchmarks, and a fourth chart, Terminal-Bench, sits lower down. The data behind the points is embedded in the page. We read it from there and checked the rendered Terminal-Bench chart against it. Dividing each Haiku 5.5 point's cost by Haiku 4.5's single plotted point gives the table below.
Cost per attempt and score for Haiku 5.5 at its default and at Max, against Haiku 4.5's one plotted point. Chart data embedded in the launch page; the comparisons are our division
- Benchmark, and Haiku 4.5's point
- OSWorld 2.1 partial credit: 15.7% at $1.45
- Medium (the default)
- 53.3% at $0.126, 91% lower
- Max (the table's setting)
- 72.4% at $0.611, 58% lower
- Benchmark, and Haiku 4.5's point
- GDPval-AA v2.1: 735 Elo at $0.24
- Medium (the default)
- 1277 at $0.030, 87% lower
- Max (the table's setting)
- 1620 at $0.866, 3.6 times as much
- Benchmark, and Haiku 4.5's point
- Humanity's Last Exam, no tools: 10.2% at $0.093
- Medium (the default)
- 35.5% at $0.0069, 93% lower
- Max (the table's setting)
- 45.9% at $0.214, 2.3 times as much
- Benchmark, and Haiku 4.5's point
- Terminal-Bench 4.0: 0.0% at $0.79
- Medium (the default)
- 20.3% at $0.680, 14% lower
- Max (the table's setting)
- 39.2% at $2.64, 3.3 times as much
| Benchmark, and Haiku 4.5's point | Medium (the default) | Max (the table's setting) |
|---|---|---|
| OSWorld 2.1 partial credit: 15.7% at $1.45 | 53.3% at $0.126, 91% lower | 72.4% at $0.611, 58% lower |
| GDPval-AA v2.1: 735 Elo at $0.24 | 1277 at $0.030, 87% lower | 1620 at $0.866, 3.6 times as much |
| Humanity's Last Exam, no tools: 10.2% at $0.093 | 35.5% at $0.0069, 93% lower | 45.9% at $0.214, 2.3 times as much |
| Terminal-Bench 4.0: 0.0% at $0.79 | 20.3% at $0.680, 14% lower | 39.2% at $2.64, 3.3 times as much |
Artificial Analysis, which runs its own tests, shows the same shape. On its Intelligence Index, read from its leaderboard on 8 October, Haiku 5.5's printed cost per task is $0.02 at Low, $0.05 at Medium, $0.08 at High, $0.12 at Xhigh and $0.21 at Max, against $0.28 for Haiku 4.5 with reasoning. That is 93, 82, 71, 57 and 25 per cent lower (derived; Artificial Analysis rounds to the cent), while the index score rises from Haiku 4.5's 17 to 29, 34, 38, 41 and 43. Anthropic's 75 per cent matches the Medium to High range on both sets of numbers. It does not describe Max.
Cheaper per token is not cheaper per task: the Terminal-Bench chart
The launch page's last chart, in its section on Sonnet's cache price, matters most to anyone choosing between the two models. Terminal-Bench 4.0 tests complex terminal work. The page says Sonnet 5.5 and Opus 5.5 "remain better choices" for it and that Haiku 5.5 suits "more narrowly scoped tasks". The chart data say more than that.
At Medium both models cost $0.68 an attempt, and Haiku scores 20.3 per cent against Sonnet's 28.8. At Low, Sonnet costs $0.62 for 20.0 per cent, less than Haiku's Medium for the same score. Haiku at Max costs $2.64 for 39.2 per cent, while Sonnet at High costs $1.46 for 43.0. For every Haiku point except Low, a Sonnet point costs no more and scores higher (derived). Because Haiku's input and output price is a twentieth of Sonnet's, Haiku at Max, at 1.8 times Sonnet High's cost, implies it used several times to a few dozen times the tokens: 3.6 to 36 times, depending on how much of the traffic was long-prompt and cache-read priced. That is our inference from the price table, not a figure Anthropic prints.
Terminal-Bench 4.0 cost per successful attempt: cost per attempt divided by pass rate. Derived from Anthropic chart data
- Effort
- Low
- Haiku 5.5
- $3.34
- Sonnet 5.5
- $3.10
- Effort
- Medium
- Haiku 5.5
- $3.35
- Sonnet 5.5
- $2.36
- Effort
- High
- Haiku 5.5
- $4.18
- Sonnet 5.5
- $3.40
- Effort
- Xhigh
- Haiku 5.5
- $5.57
- Sonnet 5.5
- $7.05
- Effort
- Max
- Haiku 5.5
- $6.74
- Sonnet 5.5
- $14.79
| Effort | Haiku 5.5 | Sonnet 5.5 |
|---|---|---|
| Low | $3.34 | $3.10 |
| Medium | $3.35 | $2.36 |
| High | $4.18 | $3.40 |
| Xhigh | $5.57 | $7.05 |
| Max | $6.74 | $14.79 |
The cheapest successful attempt on the chart is Sonnet at Medium, $2.36, against $3.34 for Haiku's cheapest. Haiku only becomes the cheaper route per pass at Xhigh and Max, where Sonnet buys far higher scores. Haiku 4.5 passed nothing, so its cost per pass is undefined.
The pattern is not universal, and the fair statement is about workloads. On OSWorld, Haiku at Xhigh scores 67.6 per cent for $0.28 against Sonnet at Medium's 66.0 for $0.93, a third of the cost. On GDPval-AA, Haiku at Xhigh and Sonnet at Medium cost the same, about $0.27, and Haiku scores 1513 against 1324. On Humanity's Last Exam, Sonnet at High scores 47.9 per cent for $0.077, above Haiku at Max's 45.9 for $0.214, 2.8 times the cost. The system card says Sonnet's costs on its Humanity's Last Exam charts assume a perfect cache-hit rate while Haiku's use recorded usage, so that last comparison is not like for like, and not in Haiku's favour.
Cognition's FrontierCode leaderboard, an outside run, points the same way for coding. It prints Haiku 5.5 at Max at 46.4 per cent, $1.33 and 181.4k output tokens a rollout. Sonnet 5.5 at Xhigh is at 52.1 per cent, $1.59 and 47.4k. Opus 5.5 at Medium is at 54.6 per cent, $0.80 and 18.5k. GPT-6 Luna at Max is at 42.4 per cent, $0.10 and 56.6k. The system card says Cognition's result files carried no cost for Haiku 5.5, and the leaderboard as read prints one. How it was added is not stated.
Every benchmark in the launch table, and who ran it
The launch table compares four models on seven benchmarks, eight rows with Humanity's Last Exam counted twice. The system card says Haiku 5.5's results use adaptive thinking at max effort, averaged over five trials unless noted, and that competitor figures come from their developers' system cards or public leaderboards unless noted.
The launch table with who ran each row and the condition. Sources: launch page; system card sections 8.1 to 8.10. Column order: Haiku 5.5 / Haiku 4.5 / GPT-6 Luna / Sonnet 5.5
- Benchmark
- GDPval-AA v2.1 (Elo)
- Scores as published
- 1620 / 735 / 1437 / 1840
- Who ran it, and the condition
- Artificial Analysis, independently. Haiku at Max. At the default Medium Haiku scores 1277
- Benchmark
- AA-Briefcase v1.1 (Elo)
- Scores as published
- 1578 / 614 / 1336 / 1824
- Who ran it, and the condition
- Artificial Analysis. Medium: 1372
- Benchmark
- OSWorld 2.1, offline subset
- Scores as published
- 72.4% / 15.7% / 48.9% / 83.9%
- Who ran it, and the condition
- Anthropic, 82 of 108 tasks. Partial credit: the strict pass rate is 37.1% for Haiku 5.5 and 48.8% for Sonnet 5.5. Anthropic ran Luna through OpenAI's API
- Benchmark
- Humanity's Last Exam, no tools, then with tools
- Scores as published
- 45.9% / 10.2% / none / 56.9%, then 57.4% / 18.7% / none / 64.5%
- Who ran it, and the condition
- Anthropic. Luna has no score
- Benchmark
- Terminal-Bench 4.0
- Scores as published
- 39.2% / 0.0% / 16.4% / 70.6%
- Who ran it, and the condition
- Anthropic, in Claude Code, no internet, safeguards on. Luna's 16.4% is the public leaderboard's
- Benchmark
- FrontierCode 1.1 Main
- Scores as published
- 46.4% / none / 42.4% / 52.1%
- Who ran it, and the condition
- Cognition. Haiku and Luna at Max, Sonnet at Xhigh. At Max Sonnet scores 46.2%
- Benchmark
- Chartography, no tools
- Scores as published
- 46.4% / 6.4% / 29.1% / 61.6%
- Who ran it, and the condition
- Anthropic's run of Surge AI's benchmark. With tools: 86.2% against Sonnet's 90.2%. Luna's figure is Surge's public one
| Benchmark | Scores as published | Who ran it, and the condition |
|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 / 735 / 1437 / 1840 | Artificial Analysis, independently. Haiku at Max. At the default Medium Haiku scores 1277 |
| AA-Briefcase v1.1 (Elo) | 1578 / 614 / 1336 / 1824 | Artificial Analysis. Medium: 1372 |
| OSWorld 2.1, offline subset | 72.4% / 15.7% / 48.9% / 83.9% | Anthropic, 82 of 108 tasks. Partial credit: the strict pass rate is 37.1% for Haiku 5.5 and 48.8% for Sonnet 5.5. Anthropic ran Luna through OpenAI's API |
| Humanity's Last Exam, no tools, then with tools | 45.9% / 10.2% / none / 56.9%, then 57.4% / 18.7% / none / 64.5% | Anthropic. Luna has no score |
| Terminal-Bench 4.0 | 39.2% / 0.0% / 16.4% / 70.6% | Anthropic, in Claude Code, no internet, safeguards on. Luna's 16.4% is the public leaderboard's |
| FrontierCode 1.1 Main | 46.4% / none / 42.4% / 52.1% | Cognition. Haiku and Luna at Max, Sonnet at Xhigh. At Max Sonnet scores 46.2% |
| Chartography, no tools | 46.4% / 6.4% / 29.1% / 61.6% | Anthropic's run of Surge AI's benchmark. With tools: 86.2% against Sonnet's 90.2%. Luna's figure is Surge's public one |
Who ran Luna's figures. None came from OpenAI's own claims. GDPval-AA and AA-Briefcase came from Artificial Analysis. Terminal-Bench is the public leaderboard's, and the system card says that, to Anthropic's knowledge, OpenAI has not reported its own for Luna. FrontierCode is Cognition's, Chartography is as Surge reports it, and OSWorld was run by Anthropic with OpenAI's own context compaction.
The zero. Haiku 4.5's 0.0 per cent on Terminal-Bench 4.0 is a measured result, not a missing one. It ran, at $0.79 an attempt on Anthropic's chart, with a fixed 63,999-token thinking budget in Claude Code's bare mode, and passed none of 660 trials. The system card does not say why. It defines no ratio, so "39.2 against 0.0" is a statement about that setup and not a rate of improvement. Terminal-Bench-Science gives Haiku 4.5 0.1 per cent, one success in 700 trials.
Mixed settings and measures. OSWorld's 72.4 is a partial-credit score, and the strict pass rate is 37.1 per cent. FrontierCode puts Sonnet at its best, Xhigh at 52.1, beside Haiku at Max. At the same Max, Sonnet scores 46.2 against Haiku's 46.4. Chartography's row is without tools, and with tools it closes to 86.2 against 90.2.
Safeguards on, no fallback. Haiku 5.5's Terminal-Bench run had its safeguards on and no fallback model. Of 660 trials, 12 (1.8 per cent, 10 of them on one task) were blocked, and all failed.
The customer quotes
The page says customers "reported results consistent with the performance and cost improvements shown above". It quotes six companies. None of the six prints a cost figure, and two mention cost only in words.
The numbers in the six customer quotes on the launch page, and what each leaves out. Staff names deliberately not printed
- Company
- Asana
- Figure printed
- Over 30% lower latency and up to 2.5 times faster inference per agent turn, on its own agent evals
- Not printed
- The model it compared against, described only as the one it uses today
- Company
- HubSpot
- Figure printed
- 92.8% averaged over three runs on its CRM suite, the best it has seen
- Not printed
- Which models, the other scores, the cost
- Company
- AlphaSense
- Figure printed
- 0.84 against 0.76 for Haiku 4.5 over 400 queries, for a feature it says makes about 8M calls a week
- Not printed
- Cost, and what the score measures
- Company
- Box
- Figure printed
- 11 points above Haiku 4.5 at about half the latency
- Not printed
- The task set and the scale
- Company
- Rogo
- Figure printed
- None
- Not printed
- Any figure
- Company
- Cognition
- Figure printed
- A FrontierCode score of 66.2 for Devin Fusion with Haiku 5.5 as the sidekick and Opus 5.5 as the lead
- Not printed
- Haiku 5.5 alone scores 46.4% on FrontierCode Main. The pairing's cost is not printed
| Company | Figure printed | Not printed |
|---|---|---|
| Asana | Over 30% lower latency and up to 2.5 times faster inference per agent turn, on its own agent evals | The model it compared against, described only as the one it uses today |
| HubSpot | 92.8% averaged over three runs on its CRM suite, the best it has seen | Which models, the other scores, the cost |
| AlphaSense | 0.84 against 0.76 for Haiku 4.5 over 400 queries, for a feature it says makes about 8M calls a week | Cost, and what the score measures |
| Box | 11 points above Haiku 4.5 at about half the latency | The task set and the scale |
| Rogo | None | Any figure |
| Cognition | A FrontierCode score of 66.2 for Devin Fusion with Haiku 5.5 as the sidekick and Opus 5.5 as the lead | Haiku 5.5 alone scores 46.4% on FrontierCode Main. The pairing's cost is not printed |
These are early-access testimonials selected by the vendor, which says nothing against the good faith of the companies named. They are the only workload-level evidence on the page, and they carry no method.
What independent trackers print
Anthropic sells Haiku 5.5, chose the benchmarks and the rival, and ran most of the evaluations itself. That is normal practice, and it is why the outside rows matter. Six public sources carried some Haiku 5.5 data, or the lack of it, on 8 October.
Independent and public sources for Haiku 5.5, read on 8 October 2026 between about 13:40 and 14:00 BST. Live pages, so figures can change
- Source
- Artificial Analysis, GDPval-AA and AA-Briefcase
- What it prints
- Haiku 5.5 at Max 1620 and 1578, at Medium 1277 and 1372. Haiku 4.5 735 and 614. GPT-6 Luna at Max 1432 and 1336
- Against Anthropic's table
- Agrees within its intervals. Luna's GDPval-AA reads 1432 against Anthropic's 1437
- Source
- Artificial Analysis, Intelligence Index v4.3.2
- What it prints
- Haiku 5.5 at Max 43 for $0.21 a task. Sonnet 5.5 at Medium 41 for $0.48. GPT-6 Luna at Max 38 for $0.07. Haiku 4.5 17 for $0.28
- Against Anthropic's table
- Not in Anthropic's table. Haiku is the cheaper route against Sonnet, and level with Luna
- Source
- Vals AI
- What it prints
- Vals Index 54.31%, rank 16 of 45, at $2.99 a test, with Haiku at compute effort Max. Terminal-Bench 4.0 35.35%, plus or minus 2.20. Sonnet 5.5 67.04% at $21.34. Luna 51.22% at $0.43
- Against Anthropic's table
- Terminal-Bench is 3.85 points below Anthropic's 39.2. Vals puts Sonnet 5.5 at 64.14 (Anthropic 70.6) and Luna at 13.64 (16.4)
- Source
- Terminal-Bench public leaderboard
- What it prints
- No Haiku 5.5 or Haiku 4.5 row. Sonnet 5.5 at Max 61.8%, plus or minus 2.9. GPT-6 Luna at Max 16.4%
- Against Anthropic's table
- Luna agrees. Sonnet is 8.8 points below Anthropic's 70.6
- Source
- Cognition, FrontierCode
- What it prints
- Haiku 5.5 at Max 46.4% on Main, $1.33 and 181.4k output tokens a rollout
- Against Anthropic's table
- Agrees on the score
- Source
- OSWorld 2.0 project page and Surge Chartography
- What it prints
- No Haiku 5.5 or Sonnet 5.5 row on either. Surge lists GPT-6 Luna at 29.1%
- Against Anthropic's table
- Luna's figure agrees. Haiku's Chartography and OSWorld scores rest on Anthropic's own run
| Source | What it prints | Against Anthropic's table |
|---|---|---|
| Artificial Analysis, GDPval-AA and AA-Briefcase | Haiku 5.5 at Max 1620 and 1578, at Medium 1277 and 1372. Haiku 4.5 735 and 614. GPT-6 Luna at Max 1432 and 1336 | Agrees within its intervals. Luna's GDPval-AA reads 1432 against Anthropic's 1437 |
| Artificial Analysis, Intelligence Index v4.3.2 | Haiku 5.5 at Max 43 for $0.21 a task. Sonnet 5.5 at Medium 41 for $0.48. GPT-6 Luna at Max 38 for $0.07. Haiku 4.5 17 for $0.28 | Not in Anthropic's table. Haiku is the cheaper route against Sonnet, and level with Luna |
| Vals AI | Vals Index 54.31%, rank 16 of 45, at $2.99 a test, with Haiku at compute effort Max. Terminal-Bench 4.0 35.35%, plus or minus 2.20. Sonnet 5.5 67.04% at $21.34. Luna 51.22% at $0.43 | Terminal-Bench is 3.85 points below Anthropic's 39.2. Vals puts Sonnet 5.5 at 64.14 (Anthropic 70.6) and Luna at 13.64 (16.4) |
| Terminal-Bench public leaderboard | No Haiku 5.5 or Haiku 4.5 row. Sonnet 5.5 at Max 61.8%, plus or minus 2.9. GPT-6 Luna at Max 16.4% | Luna agrees. Sonnet is 8.8 points below Anthropic's 70.6 |
| Cognition, FrontierCode | Haiku 5.5 at Max 46.4% on Main, $1.33 and 181.4k output tokens a rollout | Agrees on the score |
| OSWorld 2.0 project page and Surge Chartography | No Haiku 5.5 or Sonnet 5.5 row on either. Surge lists GPT-6 Luna at 29.1% | Luna's figure agrees. Haiku's Chartography and OSWorld scores rest on Anthropic's own run |
Vals AI also prints a fallback rate and a refusal rate. Haiku 5.5 shows 0.00 per cent fallback and 0.22 per cent refusals on its suite, against 4.56 and 0.13 per cent for Sonnet 5.5, and 0.04 per cent refusals for Luna. The harnesses differ, so levels differ: Artificial Analysis's own Terminal-Bench 4.0 page headlines Sonnet 5.5 at 63.6 per cent. Only comparisons within one source are like for like. On the three sources that print a cost for Haiku 5.5 (Artificial Analysis, Vals AI and Cognition), Sonnet 5.5 at the settings listed costs more than Haiku 5.5, and Haiku 5.5 at Max costs more than GPT-6 Luna at Max.
Fast per token, slower to answer
Footnote 1 calls Haiku 5.5 Anthropic's fastest model to date "at each model's standard speed". The page prints no speed figure, and the customer quotes give latency changes against unnamed baselines. Artificial Analysis measures about 242 output tokens a second for Haiku 5.5 at Max, against 90 for Haiku 4.5, 101 to 129 for Sonnet 5.5 and 112 to 129 for GPT-6 Luna at the levels it lists.
Its end-to-end figure, seconds to receive a 500-token answer including thinking, tells a different story at matched effort. Haiku 5.5 takes 12.7 seconds at Low and 17.1 at Medium, against 6.0 and 7.0 for Sonnet 5.5, and 19.8 for Haiku 4.5. At Max it takes 434 seconds, almost all of it before the first token, which is thinking time. Artificial Analysis's model page and leaderboard print different first-token latencies, 323 and 432 seconds, so treat those as orders of magnitude. Anthropic's system card shows the same shape on its own tests: 8 to 21 seconds per answer on HealthBench Professional from Low to Xhigh, and about 110 seconds at Max. Anthropic's docs recommend Low for chat and short tool tasks, and warn that at Low the model is more likely to skip a search, stop early or skip a check.
"Fastest" holds for token output. Whether a live-support or browser workload is fastest depends on the effort level you pay for.
For security teams: narrower blocks, no fallback, and a model not yet built for the attacker's job
Anthropic says Haiku 5.5's cyber safeguards are more restrictive than Haiku 4.5's and somewhat less restrictive than those on other recent models. They permit a wider range of defensive tasks than Sonnet 5.5's safeguards but "still block penetration testing", and techniques more likely to be used by attackers.
The system card and docs add the mechanism. The classifiers have no fallback model on Anthropic's own products or its API. A blocked request ends in a response with stop_reason set to refusal, and the docs say that re-sending it "usually returns another refusal". The system card adds that traffic through other platforms "may experience different behavior". Benign security work can trigger the cyber category, and the docs say that for anyone moving from Haiku 4.5 these refusals are new. Vals AI recorded a 0.22 per cent refusal rate across its suite.
Cyber capability with safeguards switched off. Source: Haiku 5.5 system card, section 3.3; Sonnet 5.5 ExploitBench from our Sonnet briefing
- Evaluation
- CyScenarioBench, 10 multi-stage challenges built by Irregular
- Haiku 5.5
- 3.3% solved
- Sonnet 5.5
- 46.1% solved
- Evaluation
- Binary Exploitation Benchmark, control-flow hijacks from 831 entry points
- Haiku 5.5
- 3
- Sonnet 5.5
- 50
- Evaluation
- ExploitBench, full arbitrary code execution across 410 runs
- Haiku 5.5
- 4 (1.0%)
- Sonnet 5.5
- 178 (43.4%)
| Evaluation | Haiku 5.5 | Sonnet 5.5 |
|---|---|---|
| CyScenarioBench, 10 multi-stage challenges built by Irregular | 3.3% solved | 46.1% solved |
| Binary Exploitation Benchmark, control-flow hijacks from 831 entry points | 3 | 50 |
| ExploitBench, full arbitrary code execution across 410 runs | 4 (1.0%) | 178 (43.4%) |
That is a model that cannot yet do the attacker's job at any scale, which is why its blocks are narrower. It is not the same as safe to leave unwatched, because prompt injection is the relevant risk for a subagent that reads untrusted text. On Gray Swan's indirect prompt injection benchmark, an attacker succeeded against Haiku 5.5 7.1 per cent of the time within 15 attempts, against 83.2 per cent for Haiku 4.5, 3.4 for Sonnet 5.5 and 1.0 for Opus 5.5. In GUI computer use the Haiku 5.5 figure is 24.4 per cent. In Shade's coding attacks 0.08 per cent of attempts succeeded, and because Haiku 5.5 has no fallback, every answer came from Haiku itself. The Sonnet 5.5 figure in our Sonnet briefing included requests re-routed to an older model.
The alignment section reports three things worth reading before giving the model a job. It over-refused more than any other model in the automated behavioural audit, while refusing fewer benign single-turn requests than Haiku 4.5 (0.17 against 0.44 per cent on the API). It used a leaked answer without telling the user 17 per cent of the time, up from 2 per cent for Haiku 4.5. And it hallucinated about as much as Haiku 4.5.
For security teams blocked on legitimate work, Anthropic's pages point to the Cyber Verification Program. The program page published on 6 October names Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1 and "new models moving forward". It does not name Haiku 5.5, and the Help Center article does not mention Haiku at all. Whether a verification tier lowers Haiku 5.5's blocks is not stated on any page read. Enrolled organisations must accept data retention so that Anthropic can monitor for cyber misuse. Our briefing on the program covers the tiers and the 34 of 50 against 4 of 50 result.
A customer on Haiku 4.5 has a decision, but not yet a date
Anthropic's deprecations page lists claude-haiku-4-5-20251001 as Active, with a tentative retirement "not sooner than October 15, 2026". That is seven days after the date of this briefing (derived), and no deprecation notice is listed. The page promises at least 60 days' notice before the retirement of a publicly released model, so a notice on 8 October would put the earliest retirement at 7 December 2026 (derived). The precedent is quick. Sonnet 4.5's notice came on 30 September, two days after Sonnet 5.5 launched, with retirement on 30 November, 61 days later (derived). The Haiku 5.5 migration guide already tells Haiku 4.5 customers to check the deprecations page. Amazon Bedrock and Google Cloud set their own retirement schedules.
The decision is whether to migrate on your schedule or on Anthropic's, and the move is not a model ID swap. The migration guide lists: manual thinking budgets now return an error and adaptive thinking is on by default; temperature, top_p and top_k are removed; assistant prefill is rejected; computer use moves to a new toolset; thinking blocks are tied to the account that produced them; a refusal stop reason is new; the same text counts as about 30 per cent more tokens; and Priority Tier is not supported. Each is a change to test, and the migration can raise spend before it lowers it.
For UK organisations: treat the swap as a change, and test it as a model
For a UK organisation that uses Haiku 5.5 as a subagent, summariser or classifier, the cost story is the smaller half. ISO/IEC 42001:2023, which ISO describes as specifying requirements for establishing, implementing, maintaining and continually improving an AI management system, is the obvious frame for the inventory and change control below. We read ISO's catalogue page, not the paid standard, so this briefing makes no claim about its clauses.
The NCSC supplies the rest. Its August 2026 advice on agentic AI says a model's built-in safeguards "should not be treated as holistic", that judge models "should be independently evaluated", and that named individuals or groups should be responsible. Its September 2026 post on defending with agents rates an AI action by how much it can change, how far it reaches, how critical the target is, how well it can be proved in advance and how easily it can be reversed, and suggests starting with tasks that "advise a human". Its December 2025 post says current models "do not enforce a security boundary" between instructions and data, so the aim is to reduce likelihood and impact.
For this model, an inventory entry needs the ID, since claude-haiku-5-5 has no snapshot suffix and the ID is the only pin. It needs the platform and region the model is called through, because on Bedrock and Google Cloud the cloud provider is the data processor and on the Claude API, Claude Platform on AWS and Foundry Anthropic is. It needs the effort setting and whether thinking is on, the prompt and workflow version, what the model may decide, the owner and the date of the last evaluation. A swap from Haiku 4.5 changes the tokenizer, the refusal behaviour, the thinking behaviour and the retention and residency options. It is a change under your change process, not a configuration edit, and your own evaluation has to be re-run before it ships.
For UK organisations: where prompts go and how long they stay
What Anthropic's pages state about data handling for Haiku 5.5 users, and what they do not. Sources: Commercial Terms, privacy centre retention article, API and data retention and data residency documentation, regional compliance page, cloud platform documentation, read 8 October 2026
- Topic
- Contract and law
- Stated on the pages read
- Customers in the EEA, Switzerland or the UK contract with Anthropic Ireland, Limited, under Irish law. Anthropic 'may not train models on Customer Content from Services' (Commercial Terms, effective 17 June 2025 as read)
- Not stated
- Whether any UK-specific terms apply. The data processing addendum was not read
- Topic
- Retention, Anthropic API
- Stated on the pages read
- Inputs and outputs deleted from Anthropic's back end within 30 days, except where a service keeps them longer, an agreement says otherwise, the law requires it or the usage policy has to be enforced (privacy centre, dated 1 July 2026)
- Not stated
- The API docs say conversation content is "not retained by default". The two pages word the default differently. The contract decides
- Topic
- Flagged content
- Stated on the pages read
- If flagged by automated trust and safety systems: inputs and outputs for up to 2 years, classification scores up to 7 years. The docs say up to 2 years even under zero data retention
- Not stated
- Whether a cyber-classifier block on benign work counts as flagged. The docs say these refusals are new since Haiku 4.5
- Topic
- Zero data retention
- Stated on the pages read
- On request through sales, per organisation. The 30-day retention requirement covers Fable and Mythos models, and Haiku 5.5 is not among them
- Not stated
- Whether it applies under the Cyber Verification Program for Haiku 5.5, which requires retention for enrolled organisations
- Topic
- Residency, Claude API
- Stated on the pages read
- inference_geo takes global (the default) or us, at 1.1 times the price. Workspace geo is us only. Haiku 4.5 rejects the parameter with a 400 error
- Not stated
- Any UK or EU value. Haiku 5.5 is not named among supported models: the page says Claude 4.6 and later
- Topic
- Residency, cloud platforms
- Stated on the pages read
- Bedrock lists Europe (London), eu-west-2, among its regions and offers an EU inference profile. Google Cloud offers us and eu multi-region endpoints. Foundry hosts Haiku 5.5 on Azure or on Anthropic. Hosted on Azure, prompts and completions stay in Azure and only usage metadata and flagged content go to Anthropic
- Not stated
- Which models each Bedrock region serves. The regional compliance page lists Europe, not the UK, shows Foundry's Europe as coming soon, and names Haiku 4.5 rather than Haiku 5.5
| Topic | Stated on the pages read | Not stated |
|---|---|---|
| Contract and law | Customers in the EEA, Switzerland or the UK contract with Anthropic Ireland, Limited, under Irish law. Anthropic 'may not train models on Customer Content from Services' (Commercial Terms, effective 17 June 2025 as read) | Whether any UK-specific terms apply. The data processing addendum was not read |
| Retention, Anthropic API | Inputs and outputs deleted from Anthropic's back end within 30 days, except where a service keeps them longer, an agreement says otherwise, the law requires it or the usage policy has to be enforced (privacy centre, dated 1 July 2026) | The API docs say conversation content is "not retained by default". The two pages word the default differently. The contract decides |
| Flagged content | If flagged by automated trust and safety systems: inputs and outputs for up to 2 years, classification scores up to 7 years. The docs say up to 2 years even under zero data retention | Whether a cyber-classifier block on benign work counts as flagged. The docs say these refusals are new since Haiku 4.5 |
| Zero data retention | On request through sales, per organisation. The 30-day retention requirement covers Fable and Mythos models, and Haiku 5.5 is not among them | Whether it applies under the Cyber Verification Program for Haiku 5.5, which requires retention for enrolled organisations |
| Residency, Claude API | inference_geo takes global (the default) or us, at 1.1 times the price. Workspace geo is us only. Haiku 4.5 rejects the parameter with a 400 error | Any UK or EU value. Haiku 5.5 is not named among supported models: the page says Claude 4.6 and later |
| Residency, cloud platforms | Bedrock lists Europe (London), eu-west-2, among its regions and offers an EU inference profile. Google Cloud offers us and eu multi-region endpoints. Foundry hosts Haiku 5.5 on Azure or on Anthropic. Hosted on Azure, prompts and completions stay in Azure and only usage metadata and flagged content go to Anthropic | Which models each Bedrock region serves. The regional compliance page lists Europe, not the UK, shows Foundry's Europe as coming soon, and names Haiku 4.5 rather than Haiku 5.5 |
The practical points follow from the table. On the pages read, a UK organisation that needs prompts to stay in the UK or the EU has no first-party option and must look to a cloud platform, where the processor is the cloud provider and the retention terms are that provider's. Platforms also differ in features: Bedrock does not support structured outputs, the Batch API or server-side tools for Claude. And the new classifiers raise one question. Hosted on Azure, content flagged by Anthropic's safety systems goes to Anthropic, and on the first-party API flagged content is kept for up to 2 years. For security teams, a request that the cyber classifier blocks may therefore be the one that leaves your cloud boundary. That is our inference, because the pages do not say that a block counts as a flag.
For UK organisations: measure cost per task, not per token
Price is per token and spend is per task. Cost per task is every input, cached and output token a task consumes, at the price for that request's prompt length, across retries and tool calls, divided by the tasks that succeeded. Anthropic's own prompting guide tells Haiku users to compare two or three effort levels on their own evals, to use Xhigh or Max only where a quality gain justifies the cost, and to run the same evals on Sonnet 5.5. The cost-per-pass table above is that exercise on Anthropic's data.
Record three more numbers beside each score. The first is the effort level: Haiku 5.5 defaults to Medium, and the launch table's scores are at Max. The second is the share of requests over 100,000 tokens, which pay up to five times the price. The third is output tokens per task, because Artificial Analysis found Haiku 5.5 verbose at Max. Set spend alerts on each workspace or key, and review them after any change of effort or prompt. A change of effort alone moved Haiku 5.5's cost per attempt by 3.9 to 31 times between Medium and Max on Anthropic's charts (derived).
What to do, in the order worth doing
Take this with you
A defender checklist for adopting Haiku 5.5 as a subagent or classifier
- Inventory where a Haiku-class model sits in your workflows today: model ID, platform and region, effort setting, owner, and what it is allowed to decide. Count Haiku 4.5 separately, because it has a retirement date coming.
- Run your own evaluation of about 50 tasks from your own data at Low, Medium and High, and at Max only if the gain justifies it. Record cost per task, cost per successful task and the type of each error, for Haiku 5.5, Haiku 4.5 and Sonnet 5.5.
- Pin claude-haiku-5-5 as the model ID, log the model and usage fields of every response, and record any change of effort, prompt or platform as a change under your change process.
- Set spend alerts on each workspace or key, and review them after any change of effort or prompt.
- Decide what the cheap model may never decide alone: closing an alert, changing a system, granting access, or acting on text it read from an untrusted source. Use the NCSC idea of potency: advise a human first.
- Test your own security use cases for refusals. There is no fallback, so handle stop_reason refusal in your client and decide what happens to the work. If blocks stop legitimate security work, read the Cyber Verification Program terms, including data retention, before applying.
- Check the data position on the platform you actually use: contracting entity, retention, flagged content, residency, and who the data processor is. Do not assume a UK or EU option for Haiku 5.5 that the pages do not state.
- Plan for the Haiku 4.5 retirement: watch the deprecations page, find any Priority Tier commitment (not supported on Haiku 5.5), and budget a migration test that covers the tokenizer, thinking and refusal changes.
The question this launch leaves
Haiku 5.5 may well be the cheapest way to do a great deal of summarising, routing and classification, and on OSWorld and GDPval-AA Anthropic's own data say it is. The headline is an average at a setting the page does not name, and the benchmark table is at a setting most customers will not run. The question is for the buyer.
At which effort setting does your cheapest workflow run today, who chose it, and would anyone notice if it changed?
Key facts
Sources
- PrimaryIntroducing Claude Haiku 5.5, 7 October 2026, read in full by HTTP and again in a browser tab: the benchmark table, the price table, footnotes 1 and 2, safety and availability text, the customer quotes, and the data behind the four accuracy against cost charts, read from the page and checked against the rendered Terminal-Bench chartAnthropicaccessed 2026-10-08
- PrimarySystem Card: Claude Haiku 5.5, 7 October 2026, 144 pages read as text: safeguards and the absence of a fallback, cyber evaluations, prompt injection, alignment findings, and the capability methods, including who ran each benchmark and the effort and response-time resultsAnthropicaccessed 2026-10-08
- PrimaryClaude Haiku 5.5 migration guide: the tokenizer change of about 30% more tokens, the fixed model ID, thinking, sampling, prefill and computer use changes, the refusal stop reason and the lack of Priority Tier supportAnthropicaccessed 2026-10-08
- PrimaryModels overview: Haiku 5.5 default effort, context window, retirement date and model IDs on each platformAnthropicaccessed 2026-10-08
- PrimaryModel deprecations: Haiku 4.5 listed Active with a retirement not sooner than 15 October 2026, the 60 days notice commitment, and the Sonnet 4.5 notice and retirement datesAnthropicaccessed 2026-10-08
- PrimaryPricing documentation: Haiku 5.5 priced by prompt length, cache write rates, batch prices and the 1.1 times multiplier for US-only inferenceAnthropicaccessed 2026-10-08
- PrimaryPrompting Claude Haiku 5.5: effort guidance, thinking on by default, the doubling of output tokens from Low to Medium, and the safeguard refusal categories with no server-side fallbackAnthropicaccessed 2026-10-08
- PrimaryEffort documentation: the five levels, Medium as the Haiku 5.5 default, and the recommendation to compare Xhigh and Max with Sonnet 5.5Anthropicaccessed 2026-10-08
- PrimaryData residency documentation: inference_geo values global and us, workspace geo us only, and the 400 error on Haiku 4.5Anthropicaccessed 2026-10-08
- PrimaryAPI and data retention documentation: zero data retention on request, Covered Models, the 2 year retention of flagged content, and which platforms have Anthropic or the cloud provider as processorAnthropicaccessed 2026-10-08
- PrimaryPrivacy centre article on commercial data retention, dated 1 July 2026: 30 day deletion of API inputs and outputs, and 2 year and 7 year retention when flaggedAnthropicaccessed 2026-10-08
- PrimaryCommercial Terms of Service, effective 17 June 2025 as read: Anthropic Ireland, Limited for customers in the EEA, Switzerland or the UK, Irish law, and the commitment not to train on customer contentAnthropicaccessed 2026-10-08
- PrimaryClaude in Amazon Bedrock documentation: Haiku 5.5 access criteria, the region table including Europe (London), and unsupported featuresAnthropicaccessed 2026-10-08
- PrimaryClaude in Microsoft Foundry documentation: Haiku 5.5 hosted on Azure or on Anthropic, and what leaves AzureAnthropicaccessed 2026-10-08
- PrimaryRegional compliance page: regional availability by platform, the Europe row and the model names in its FAQAnthropicaccessed 2026-10-08
- PrimaryHaiku product page: availability, pricing text and use casesAnthropicaccessed 2026-10-08
- PrimaryExpanding the Cyber Verification Program, 6 October 2026: the models named, the data retention requirement and the platformsAnthropicaccessed 2026-10-08
- PrimaryLLM leaderboard read on 8 October 2026: Intelligence Index v4.3.2, cost per task, output speed and response time for Haiku 5.5 at each effort level, Haiku 4.5, Sonnet 5.5 and GPT-6 LunaArtificial Analysisaccessed 2026-10-08
- PrimaryClaude Haiku 5.5 (Max) model page: 440M output tokens on the Intelligence Index against a median of 100M, and the speed and latency textArtificial Analysisaccessed 2026-10-08
- PrimaryGDPval-AA v2.1 leaderboard: Haiku 5.5 at each effort level, Haiku 4.5 and GPT-6 Luna with intervalsArtificial Analysisaccessed 2026-10-08
- PrimaryAA-Briefcase v1.1 leaderboard: Haiku 5.5 at each effort level, Haiku 4.5 and GPT-6 Luna with intervalsArtificial Analysisaccessed 2026-10-08
- PrimaryTerminal-Bench 4.0 page: the headline score for Sonnet 5.5 and the mini-swe-agent harness; no Haiku 5.5 figure was readable as textArtificial Analysisaccessed 2026-10-08
- PrimaryClaude Haiku 5.5 on Vals AI, read in a browser on 8 October 2026: Vals Index, cost per test, Terminal-Bench 4.0, refusal and fallback rates, and compute effort Max; the Sonnet 5.5 and GPT-6 Luna pages were read the same wayVals AIaccessed 2026-10-08
- PrimaryTerminal-Bench 4.0 public leaderboard, read in a browser on 8 October 2026: no Haiku 5.5 or Haiku 4.5 row; Sonnet 5.5 and GPT-6 Luna at MaxTerminal-Benchaccessed 2026-10-08
- PrimaryFrontierCode leaderboard, read in a browser on 8 October 2026: Haiku 5.5, Sonnet 5.5, Opus 5.5 and GPT-6 Luna scores, cost per rollout and output tokensCognitionaccessed 2026-10-08
- PrimaryChartography leaderboard, read in a browser on 8 October 2026: GPT-6 Luna at 29.1% and no Haiku 5.5 or Sonnet 5.5 entrySurge AIaccessed 2026-10-08
- PrimaryOSWorld 2.0 project page, read in a browser on 8 October 2026: no Haiku 5.5 or Sonnet 5.5 entryXLANG Lab, University of Hong Kongaccessed 2026-10-08
- PrimaryIntroducing GPT-6 Sol and Luna, read in a browser on 8 October 2026: GPT-6 Luna at $0.10 input and $0.50 output per million tokensOpenAIaccessed 2026-10-08
- PrimaryManaging the cyber risk of agentic AI, 20 August 2026, read in full: built-in safeguards not holistic, independently evaluated judge models, named responsibility and oversight levelsNCSCaccessed 2026-10-08
- PrimaryOne does not simply defend agentically, 21 September 2026, read in full: the potency, scope, criticality, rollout confidence and recoverability dimensions, and starting with tasks that advise a humanNCSCaccessed 2026-10-08
- PrimaryPrompt injection is not SQL injection (it may be worse), 8 December 2025: current models enforce no security boundary between instructions and dataNCSCaccessed 2026-10-08
- PrimaryISO/IEC 42001:2023 catalogue page, read in a browser: scope as a requirements standard for an AI management system, published December 2023. The paid standard was not readISOaccessed 2026-10-08
- Reported byOur briefing on Claude Sonnet 5.5, 28 September 2026, for the earlier cache-read price, the $12.54 Terminal-Bench cost at Max, the prompt-injection fallback figure and the ExploitBench resultpk-sharma.comaccessed 2026-10-08
- Reported byOur briefing on Claude Opus 5.5, 23 September 2026, for the earlier launch in the same familypk-sharma.comaccessed 2026-10-08
- Reported byOur briefing on the Cyber Verification Program, 6 October 2026, for the tiers and the CyScenarioBench resultpk-sharma.comaccessed 2026-10-08


