Sonnet 5.5 beats Opus 5.5 on Terminal-Bench only at maximum effort, where each task costs 71 per cent more
Anthropic released Claude Sonnet 5.5 on 28 September at Sonnet 5's per-token price, with a nine-row benchmark table and a 148-page system card. The cheaper model's one win over Opus 5.5 compares two different effort settings, and its prompt-injection results turn on a fallback model.
By Parminder Kumar Sharma · · 19 min read

The one row where the cheaper model wins
Anthropic's launch table for Claude Sonnet 5.5 has one row where the new mid-price model beats its flagship. On Terminal-Bench 4.0, an agentic coding test run in a command line, Sonnet 5.5 scores 70.6 per cent and Claude Opus 5.5 scores 66.4 per cent.
The two numbers come from different settings. The system card says Sonnet's figure is at Max effort and Opus's is at Xhigh, the level below. The launch page's own cost chart prices each point. Sonnet at Max costs $12.54 a task. Opus at Xhigh costs $7.35. The cheaper model's winning run costs 71 per cent more than the flagship run it beats (12.54 divided by 7.35 is 1.71).
At Medium effort, which Anthropic says is the default in Claude Code and the Claude apps, the order reverses and the gap is not close: Opus 5.5 scores 57.6 per cent and Sonnet 5.5 scores 28.8 per cent, exactly half.
That fact does not establish that the table is misleading. Anthropic chose a consistent rule and printed it: each model's highest score, with a footnote saying Opus's comes from Xhigh because Opus scores 64.8 per cent at Max, which the system card calls "within noise". Picking each model's best is a standard way to report a benchmark. The page also says plainly, in its own words, that Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment".
It also does not establish that Sonnet is better at terminal work. The system card gives a standard error of 2.5 points for Sonnet and 2.6 for Opus, over 66 tasks run five times each. The 4.2-point gap is about 1.2 times the combined error of 3.6 points (the square root of 2.5 squared plus 2.6 squared). That is a lead, not a result.
What the numbers do establish is the thing a buyer needs to know. Sonnet 5.5 is a very large improvement on Sonnet 5 at the same per-token price, and at low and medium effort it does that job for a fraction of the old cost. It is not a cheaper way to get Opus results. At matched spend, Opus is ahead on two of the four charted benchmarks, and Anthropic's own text says the two are "comparable at a similar cost" at higher settings. Read the table as each model's ceiling, not as a menu at one price.
What Anthropic released, and what the pages leave out
The facts below are taken from the launch page, the models overview in Anthropic's documentation and the Sonnet product page, all read on 28 September 2026.
Claude Sonnet 5.5 at release. Sources: anthropic.com/claude-sonnet-5-5, the models overview and the Sonnet product page, 28 September 2026
| Item | Published | Not stated or qualified |
|---|---|---|
| Release | 28 September 2026. Second model of the Claude 5.5 family after Opus 5.5 on 22 September | Haiku 5.5 is promised "in the coming weeks" with no date |
| API model ID | claude-sonnet-5-5, also on Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS | No dated snapshot ID is listed beside the alias |
| Per-token price | $2 input, $10 output, $0.20 cache reads, $2.50 cache writes, per million tokens. Same as Sonnet 5 | The 30 per cent saving is per task, from fewer tokens, not a price cut |
| Discounts | Up to 90 per cent with prompt caching, 50 per cent with batch processing | Whether the charted costs used caching is not stated on the launch page |
| Context and output | 1M-token context window, 128K-token maximum output | The launch page itself does not mention either |
| Knowledge cutoff | Reliable cutoff and training data cutoff both June 2026 | |
| Default effort | Medium in Claude Code and the apps, High on the Claude Platform | Benchmark table figures are mostly at Max, above either default |
| Data handling | Available with zero data retention. US-only inference at 1.1 times the token price | No UK or EU residency option is mentioned |
| Retirement | Not sooner than 28 September 2027 | |
| Migration | Anyone running Sonnet with thinking off must switch to a new between_tools setting before moving | The page links a migration guide for the detail |
The price table sets Sonnet 5.5 beside Opus 5.5. Input, output and cache writes are exactly half of Opus's $4, $20 and $5. Cache reads are the same for both, at $0.20 per million tokens. For an agent that re-reads a long cached context on every turn, cache reads can be the largest line on the bill, so on that line the saving from choosing Sonnet over Opus is zero.
Every benchmark in the launch table
The table has nine rows across four models: Sonnet 5.5, Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol. Seven rows are percentages and two are Elo ratings from pairwise judging, which cannot share an axis with a percentage. They are drawn separately below, each to its own scale, with the Elo axis starting at 1,000 because an Elo rating has no meaningful zero.
The full launch table as published, with the conditions from its footnotes and the system card. Source: anthropic.com/claude-sonnet-5-5 and system card section 8
| Benchmark | Sonnet 5.5 / Sonnet 5 / Opus 5.5 / GPT-6 Sol | Condition |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% / 10.3% / 66.4% / not reported | Sonnet at Max, Opus at Xhigh. 66 tasks, run with safeguards on and no internet |
| FrontierCode 1.1, Main | 46.2% / 42.4% / 54.4% / 49.3% | Sonnet shown at Max. It scores 52.1% at Xhigh; Max ran extra review subagents that strayed out of scope |
| CursorBench 4.0 | 55.5% / 34.1% / 57.8% / not reported | Run by Cursor. Sonnet 5.5's scores came from a table Cursor sent to Anthropic |
| GDPval-AA v2.1 | 1,844 / 1,449 / 1,846 / 1,487 | Elo. Run by Artificial Analysis on a pre-release deployment with a structured-output bug |
| AA-Briefcase v1.1 | 1,811 / 1,359 / 1,822 / 1,483 | Elo. Same pre-release deployment and bug |
| Humanity's Last Exam | 64.5% / 54.9% / 67.7% / not reported | With tools. Without tools: 56.9%, 43.1%, 64.4% |
| OSWorld 2.1 | 80.1% / 57.0% / 81.8% / not reported | Partial credit. Strict pass rates: 43.5%, 25.6%, 48.7% |
| Chartography | 61.6% / 15.6% / 64.4% / 53.6% | No tools. With tools Sonnet 5.5 scores 90.2% against Opus's 89.0% |
Three patterns hold across the table. Sonnet 5.5 beats Sonnet 5 on every row, often by a very wide margin: 395 Elo points on GDPval-AA, 46 percentage points on Chartography, 23 on OSWorld. It trails Opus 5.5 on eight of nine rows, by between 1.7 and 8.2 points where the rows are percentages. And on Terminal-Bench and Chartography, Sonnet 5 scores so low (10.3 and 15.6 per cent) that the size of the jump says as much about the old model as the new one. The page does not explain why Sonnet 5 does so badly on Terminal-Bench 4.0, a version that the system card says was rebuilt to reduce sensitivity to timeouts and harnesses.
The system card's own summary table adds rows the launch page leaves out: SWE-Bench Pro 81.3 per cent against Opus's 89.9, SWE-Bench Multilingual 90.3 against 93.9, SWE-Bench Multimodal 54.3 against 61.4. Two go the other way: HealthBench Professional, 69.2 against Opus's 65.6, and AutomationBench, 44.7 against 42.5.
The cost claims check out, and they are claims against Sonnet 5
The most useful part of the launch page is the four accuracy against cost charts. Each plots every model at five effort levels, from Low to Max, against the average cost per task. The data labels on those charts carry the exact numbers, and every cost claim on the page can be checked against them. We did.
The launch page's cost claims tested against its own chart data. Our arithmetic
| Claim on the page | Chart data | Our check |
|---|---|---|
| Terminal-Bench: Medium beats Sonnet 5's best for less than a tenth of the cost | 28.8% at $0.83 against 10.3% at $11.62 | Holds. 7.1 per cent of the cost |
| FrontierCode: High matches GPT-6 Sol's best for about a fifth of the cost | 49.4% at $0.42 against 49.3% at $2.07 | Holds. 20.3 per cent |
| FrontierCode: 10 points above Sonnet 5 at High, about one fifteenth of the cost | 49.4% at $0.42 against 39.4% at $6.10 | Holds. 10.0 points, 6.9 per cent, nearer one fourteenth |
| CursorBench: Low beats Sonnet 5's best for less than a tenth | 35.8% at $0.50 against 34.1% at $7.17 | Holds. 7.0 per cent |
| AA-Briefcase: Medium beats Sonnet 5's best for about one ninth | 1,461 at $1.64 against 1,359 at $14.43 | Holds. 11.4 per cent, nearer one ninth than one eighth |
Every comparison in that table is against Sonnet 5 or GPT-6 Sol. None is against Opus 5.5 at matched spend, so we ran that comparison from the same data, picking the nearest-cost points.
Sonnet 5.5 against Opus 5.5 at the nearest matching cost per task. Our selection from the launch page chart data
| Benchmark | Sonnet 5.5 | Opus 5.5 at similar cost |
|---|---|---|
| Terminal-Bench 4.0 | 61.5% at Xhigh, $5.30 | 64.2% at High, $3.88: Opus ahead, for less |
| CursorBench 4.0 | 53.1% at Xhigh, $3.88 | 56.0% at High, $3.97: Opus ahead |
| FrontierCode 1.1 | 49.4% at High, $0.42 | 47.3% at Low, $0.40: Sonnet ahead |
| AA-Briefcase v1.1 | 1,634 at High, $3.95 | 1,642 at Medium, $4.40: level |
The pattern matches what Anthropic says in prose. Sonnet 5.5 is most distinctive at the cheap end, where Opus has no point to compare. At the expensive end the two share roughly one cost curve, and on Terminal-Bench and AA-Briefcase Sonnet's Max run is the most expensive point on the chart, at $12.54 and $29.19 a task. On AA-Briefcase that is 39 per cent more than Opus at Max ($21.05) for an Elo 11 points lower.
The friendly-name fallacy applies here. "Sonnet" reads as the economy option, and per token it is. Per task, at the settings the headline table uses, it can be the dearer of the two. Effort, not the model name, sets the bill.
Which comparison models were chosen, and who ran the tests
The launch table has one outside comparison, GPT-6 Sol, and prints it on four rows of nine. Where GPT-6 Sol has no public score, the cost charts substitute GPT-5.6 Sol, the previous generation, on Terminal-Bench and CursorBench. The page says so in a note under each chart, which is honest, but a reader glancing at the charts sees an older OpenAI model plotted beside the new Claude one.
The system card names the model that is missing. OpenAI's GPT-6 Astra scores 57.9 per cent on Terminal-Bench 4.0, below Sonnet 5.5, but 64.6 per cent on Terminal-Bench-Science against Sonnet's 59.9, and 65.5 per cent on FrontierSWE v2 against Sonnet's 61.9 and Opus's 62.3. Those rows appear in the system card and not on the launch page. Our Opus 5.5 briefing found the same thing a week ago: Opus lost two benchmarks to GPT-6 Astra. No Google or xAI model appears anywhere on the launch page.
Who produced each launch-table number. Source: system card sections 1.6 and 8
| Benchmark | Run by | What that does not establish |
|---|---|---|
| Terminal-Bench 4.0 | Anthropic, in Claude Code bare mode | Independent replication. The public leaderboard is cited only for other models |
| FrontierCode 1.1 | Cognition, the benchmark's author | Costs for Sonnet 5.5: Cognition's files carry none |
| CursorBench 4.0 | Cursor, in its production harness | Public listing: Sonnet 5.5's scores came from a table Cursor sent Anthropic, and its costs are Anthropic's estimates |
| GDPval-AA, AA-Briefcase | Artificial Analysis | A clean run: the deployment carried a bug since fixed |
| Humanity's Last Exam, OSWorld, Chartography | Anthropic | Independent replication. OSWorld uses a Claude model as grader on tasks that need one |
None of this is unusual. Vendors run most of their own evaluations, and naming the third parties who ran the rest is better practice than many launches manage. It does mean that almost every number in the launch table is one the vendor either produced or selected. The strongest claims, where an outside party ran the test in its own harness, are FrontierCode and CursorBench, and on both Sonnet 5.5 trails Opus 5.5.
The customer quotes and the worked examples
The page carries thirteen customer quotes. Several contain numbers, and they are the most concrete claims about real workloads anywhere on the page. None comes with a method.
Numerical claims in the customer quotes, as printed. Our arithmetic where shown
| Company | Claim | Not published |
|---|---|---|
| Balyasny Asset Management | 2,441 finance tasks; about 121k tokens per answer against Sonnet 5's 497k, 76 per cent fewer | The quality scores behind "scored ahead" |
| Base44 | 118 app builds, level with Opus 5; 3.6 iterations per build against 7.7, 53 per cent fewer | How builds were scored |
| Box | More accurate, 2.4 times faster, 12 per cent fewer total tokens | Task set and accuracy measure |
| Zendesk | Tickets processed 20 per cent faster than current production models | Which models, which tickets |
| Slack | Better on almost all offline evals, about 14 per cent fewer output tokens | The evals |
| Lovable | A third fewer tool calls, roughly half the shell runs | Task count |
| Unity | 90 per cent of tasks completed on its Unity Editor and coding benchmark | The benchmark or the comparison models |
| Atlassian | Rovo Agents up to 30 per cent faster than on Sonnet 5 | Median, as opposed to best case |
These are testimonials from companies that had early access and chose to be quoted, collected by the vendor. That is a method of marketing, not of measurement, and it says nothing about the good faith of any company named. A 76 per cent cut in tokens per answer at Balyasny is a striking figure and may well be true. Nobody outside Balyasny can check it.
The worked examples are the same kind of evidence. Three side-by-side animations built by Sonnet 5 and Sonnet 5.5 from one-line prompts (400 starlings in a murmuration, wind shaping sand dunes, a clock made of 24 clocks) show speed, with no token counts or timings printed. The one internal test described, a 10-slide operating review built from a public company's earnings materials, was judged ready to send by two experts. That is a single task with two judges.
For security teams: more capable, and more refusals
The system card is direct about what changed in offensive capability. Sonnet 5.5 "is not as cyber-capable" as Opus 5.5 or Claude Mythos 5.1, "but it is a significant step up" from Sonnet 5. All the capability figures below were measured with cyber safeguards switched off, which the card says reflects what verified users will get through its Cyber Verification Program.
Cyber capability evaluations, safeguards off. Source: system card section 3.2
| Evaluation | Sonnet 5.5 | Comparison |
|---|---|---|
| Binary Exploitation Benchmark: control-flow hijacks from 831 entry points in 228 OSS-Fuzz projects | 50 | Sonnet 5: 3. Mythos 5.1: 81. Opus 5.5: 106 |
| CyScenarioBench, Irregular's multi-stage operations, 10 challenges | 46.1% solved | Sonnet 5: 0.7%. Mythos 5.1: 61.7%. Opus 5.5: 67.6% |
| ExploitBench, arbitrary code execution | 178 of 410 runs, 43.4% | Cap% of 80 in the nudged arm |
Going from 3 control-flow hijacks to 50 is roughly a seventeen-fold increase. That is why Sonnet is now deployed with the three-stage cyber safeguards Anthropic uses on Opus: a probe reading the model's internal activations, a lightweight classifier running on the model itself, and a separate trained classifier model that decides with the probe whether to block. The policy is the same as Opus's. Finding vulnerabilities in source code is allowed. Finding them in compiled binaries is blocked.
The practical consequence is written into the system card in plain terms: "users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks." When a request is blocked, Sonnet 5.5 falls back to Sonnet 5. That happens automatically in Anthropic's own apps. On the API, the developer has to opt in to fallbacks, and without opting in a blocked request is simply blocked. Traffic through other platforms "may experience different behavior".
Three other safeguards arrive with this model. Requests related to developing frontier models, such as kernel work on certain ML accelerators, are blocked and fall back to Sonnet 5. New distillation classifiers block attempts to extract hidden reasoning, with no fallback. And thinking is now tied to the account that produced it, so switching accounts in the middle of a Claude Code session changes behaviour. Biology safeguards are unchanged from Sonnet 5.
The fallback is where the prompt injections landed
Anthropic calls Sonnet 5.5 its "most robust Sonnet-class model to date" against prompt injection, and the system card supports that. The detail inside one of its tables is the most useful finding in the release for anyone deploying a coding agent.
Gray Swan's Shade tool attacked the model in 40 coding scenarios, 200 attempts each. Overall, 3.01 per cent of attempts succeeded against Sonnet 5.5 as deployed, against 19.47 per cent for Sonnet 5. But many of the injected instructions ask for destructive actions such as wiping a disk, which trips the cyber classifier and re-routes the request to Sonnet 5. That happened to 25 per cent of requests, and 12.01 per cent of those re-routed requests were compromised. Sonnet 5.5 itself was compromised in 4 of the 5,901 requests it answered, 0.07 per cent.
Weighting the two parts reproduces the headline: a quarter at 12.01 per cent plus three quarters at 0.07 per cent is 3.05 per cent, against the published 3.01. On our arithmetic, about 1,970 requests went to the fallback and roughly 236 of them were compromised, against 4 compromises of the new model. That is about 98 per cent of the successful attacks. These are derived figures, not ones Anthropic prints, and they assume the share of requests re-routed is close to exactly a quarter.
The finding has limits, and they matter. It is one benchmark in one setting. On Gray Swan's separate indirect-injection benchmark, 11 per cent of Sonnet 5.5's rollouts fell back and the fallback did not make things worse: 1 success in 777 fallback rollouts against 18 in 6,347 others. In computer use, under 1 per cent of requests fell back and neither successful attack came from the fallback. In Anthropic's browser test run through Claude Cowork, no attack succeeded in 110 scenarios, even without safeguards, a first for any model it has tested. The Shade attacker was also trained against earlier Claude models before being run unchanged, and it is our inference, not a finding, that attacks tuned that way may transfer best to an older model.
The mechanism is still worth understanding. A safety control that re-routes a suspicious request to a weaker model can hand an attacker the weaker model, and an injection that looks destructive is exactly the kind that trips the control. Anthropic reported the split rather than only the flattering total, and without it nobody outside the company could have seen this.
Two refusal measures moved the other way. On 61 malicious Claude Code prompts run ten times each, Sonnet 5.5 refused 85.2 per cent against Sonnet 5's 87.9. On 112 malicious computer-use tasks it refused 79.46 per cent against 84.68, the same rate as Opus 5.5. Both figures are without production safeguards, which add the classifiers described above.
What the alignment section says
Anthropic describes Sonnet 5.5 as not advancing the frontier of its models' capabilities, and its Responsible Scaling Policy evaluations found it crosses no new threshold. The alignment work therefore targeted risks that apply at any capability level. On an automated audit of roughly 1,850 scenarios it matches or improves on Sonnet 5 on most measures, and it has the lowest rate of attempted sandbox escapes of any model tested.
Two cautions sit in the same pages. The system card says Sonnet 5.5's "thinking is more illegible than many previous models", which matters because readable reasoning is one of the ways misbehaviour gets caught. And the launch page repeats a line worth quoting in any risk register: Sonnet 5.5 "may have tendencies we haven't found". The system card is also shorter than earlier ones by design: Anthropic says it will condense cards for non-frontier models and omit evaluations that need a lot of human time when they are not critical.
What to check before switching
Take this with you
In the order worth doing
- Decide the effort level first and test at it. The launch table's scores are mostly at Max, above both the Medium default in Claude Code and the High default on the Claude Platform.
- Price the switch per task on your own workload. The per-token price is unchanged from Sonnet 5, and cache reads cost the same as on Opus 5.5.
- If your team does security work, test your real prompts before migrating. The system card warns of more refusals even on benign cybersecurity tasks.
- On the API, decide whether to opt in to fallbacks. Without it a blocked request fails; with it the request is answered by Sonnet 5.
- Log the model that actually served each response, not only the one requested, so fallback traffic is visible.
- Treat destructive tool calls from any coding agent as needing approval. The injection results show the fallback path is where attacks succeeded.
- If you run Sonnet with thinking off, move to the between_tools setting before switching, as the migration guide requires.
- Check whether switching accounts mid-session is part of any workflow, because preserved thinking is now tied to the account.
- Keep Opus 5.5 for open-ended work. Anthropic's own text says it remains clearly stronger there.
The question this launch raises
Sonnet 5.5 is a well-documented release. The benchmarks are all printed, the conditions are in the footnotes, the cost data sits behind every chart, and the one safety number that is unflattering to Anthropic's own design is broken out in its system card rather than buried in a total. Most launches give far less.
The question it leaves is for the buyer, not the vendor. If your agents run at the default effort, which model and which setting are your numbers actually coming from, and when a classifier steps in, which model answers?
Sources
- PrimaryIntroducing Claude Sonnet 5.5, 28 September 2026: benchmark table, the four accuracy against cost charts (read from their data labels), pricing, customer quotes, worked examples, safety and availability. Read in fullAnthropicaccessed 2026-09-28
- PrimarySystem Card: Claude Sonnet 5.5, 148 pages: safeguards, cyber evaluations, agentic safety and prompt injection, alignment and capability method (sections 1, 3, 5, 6 and 8 read in full)Anthropicaccessed 2026-09-28
- PrimaryModels overview: context window, maximum output, knowledge cutoff, default effort, retirement date and platform model IDsAnthropicaccessed 2026-09-28
- PrimaryClaude Sonnet product page: availability on Claude.ai, US-only inference at 1.1 times pricing, caching and batch discountsAnthropicaccessed 2026-09-28
- Reported byOur briefing on Claude Opus 5.5, 23 September 2026, for the Opus comparison and the earlier safeguard routingpk-sharma.comaccessed 2026-09-28


