P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Sonnet 5.5 beats Opus 5.5 on Terminal-Bench only at maximum effort, where each task costs 71 per cent more

Anthropic released Claude Sonnet 5.5 on 28 September at Sonnet 5's per-token price, with a nine-row benchmark table and a 148-page system card. The cheaper model's one win over Opus 5.5 compares two different effort settings, and its prompt-injection results turn on a fallback model.

By Parminder Kumar Sharma · · 19 min read

Editorial illustration for the briefing: Sonnet 5.5 beats Opus 5.5 on Terminal-Bench only at maximum effort, where each task costs 71 per cent more

The one row where the cheaper model wins

Anthropic's launch table for Claude Sonnet 5.5 has one row where the new mid-price model beats its flagship. On Terminal-Bench 4.0, an agentic coding test run in a command line, Sonnet 5.5 scores 70.6 per cent and Claude Opus 5.5 scores 66.4 per cent.

The two numbers come from different settings. The system card says Sonnet's figure is at Max effort and Opus's is at Xhigh, the level below. The launch page's own cost chart prices each point. Sonnet at Max costs $12.54 a task. Opus at Xhigh costs $7.35. The cheaper model's winning run costs 71 per cent more than the flagship run it beats (12.54 divided by 7.35 is 1.71).

At Medium effort, which Anthropic says is the default in Claude Code and the Claude apps, the order reverses and the gap is not close: Opus 5.5 scores 57.6 per cent and Sonnet 5.5 scores 28.8 per cent, exactly half.

Bar chart of Terminal-Bench 4.0 scores and cost per task at five effort levels. Low: Sonnet 5.5 20.0 per cent at 0.76 dollars, Opus 5.5 38.5 per cent at 1.29 dollars. Medium, the Claude Code and app default: Sonnet 28.8 per cent at 0.83 dollars, Opus 57.6 per cent at 2.94 dollars. High: 43.0 at 1.94 against 64.2 at 3.88. Xhigh: 61.5 at 5.30 against 66.4 at 7.35, the Opus figure in the launch table. Max: 70.6 at 12.54, the Sonnet figure in the launch table, against 64.8 at 11.24.
Terminal-Bench 4.0 at each effort level, from the data labels on Anthropic's accuracy against cost chart. The two outlined bars are the figures the launch table compares.

That fact does not establish that the table is misleading. Anthropic chose a consistent rule and printed it: each model's highest score, with a footnote saying Opus's comes from Xhigh because Opus scores 64.8 per cent at Max, which the system card calls "within noise". Picking each model's best is a standard way to report a benchmark. The page also says plainly, in its own words, that Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment".

It also does not establish that Sonnet is better at terminal work. The system card gives a standard error of 2.5 points for Sonnet and 2.6 for Opus, over 66 tasks run five times each. The 4.2-point gap is about 1.2 times the combined error of 3.6 points (the square root of 2.5 squared plus 2.6 squared). That is a lead, not a result.

What the numbers do establish is the thing a buyer needs to know. Sonnet 5.5 is a very large improvement on Sonnet 5 at the same per-token price, and at low and medium effort it does that job for a fraction of the old cost. It is not a cheaper way to get Opus results. At matched spend, Opus is ahead on two of the four charted benchmarks, and Anthropic's own text says the two are "comparable at a similar cost" at higher settings. Read the table as each model's ceiling, not as a menu at one price.

What Anthropic released, and what the pages leave out

The facts below are taken from the launch page, the models overview in Anthropic's documentation and the Sonnet product page, all read on 28 September 2026.

Claude Sonnet 5.5 at release. Sources: anthropic.com/claude-sonnet-5-5, the models overview and the Sonnet product page, 28 September 2026

ItemPublishedNot stated or qualified
Release28 September 2026. Second model of the Claude 5.5 family after Opus 5.5 on 22 SeptemberHaiku 5.5 is promised "in the coming weeks" with no date
API model IDclaude-sonnet-5-5, also on Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWSNo dated snapshot ID is listed beside the alias
Per-token price$2 input, $10 output, $0.20 cache reads, $2.50 cache writes, per million tokens. Same as Sonnet 5The 30 per cent saving is per task, from fewer tokens, not a price cut
DiscountsUp to 90 per cent with prompt caching, 50 per cent with batch processingWhether the charted costs used caching is not stated on the launch page
Context and output1M-token context window, 128K-token maximum outputThe launch page itself does not mention either
Knowledge cutoffReliable cutoff and training data cutoff both June 2026
Default effortMedium in Claude Code and the apps, High on the Claude PlatformBenchmark table figures are mostly at Max, above either default
Data handlingAvailable with zero data retention. US-only inference at 1.1 times the token priceNo UK or EU residency option is mentioned
RetirementNot sooner than 28 September 2027
MigrationAnyone running Sonnet with thinking off must switch to a new between_tools setting before movingThe page links a migration guide for the detail

The price table sets Sonnet 5.5 beside Opus 5.5. Input, output and cache writes are exactly half of Opus's $4, $20 and $5. Cache reads are the same for both, at $0.20 per million tokens. For an agent that re-reads a long cached context on every turn, cache reads can be the largest line on the bill, so on that line the saving from choosing Sonnet over Opus is zero.

Every benchmark in the launch table

The table has nine rows across four models: Sonnet 5.5, Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol. Seven rows are percentages and two are Elo ratings from pairwise judging, which cannot share an axis with a percentage. They are drawn separately below, each to its own scale, with the Elo axis starting at 1,000 because an Elo rating has no meaningful zero.

Bar chart on a 0 to 100 per cent scale of six benchmarks for Claude Sonnet 5.5, Claude Opus 5.5, Claude Sonnet 5 and GPT-6 Sol. Terminal-Bench 4.0: 70.6 at Max, 66.4 at Xhigh, 10.3, GPT-6 Sol not reported. FrontierCode 1.1 Main: 46.2 at Max with 52.1 at Xhigh, 54.4, 42.4, 49.3. CursorBench 4.0: 55.5, 57.8, 34.1, not reported. Humanity's Last Exam with tools: 64.5, 67.7, 54.9, not reported. OSWorld 2.1 partial: 80.1, 81.8, 57.0, not reported. Chartography without tools: 61.6, 64.4, 15.6, 53.6.
The percentage rows of Anthropic's launch table, drawn from zero on one scale. GPT-6 Sol appears only where Anthropic's table gives it a score.

The full launch table as published, with the conditions from its footnotes and the system card. Source: anthropic.com/claude-sonnet-5-5 and system card section 8

BenchmarkSonnet 5.5 / Sonnet 5 / Opus 5.5 / GPT-6 SolCondition
Terminal-Bench 4.070.6% / 10.3% / 66.4% / not reportedSonnet at Max, Opus at Xhigh. 66 tasks, run with safeguards on and no internet
FrontierCode 1.1, Main46.2% / 42.4% / 54.4% / 49.3%Sonnet shown at Max. It scores 52.1% at Xhigh; Max ran extra review subagents that strayed out of scope
CursorBench 4.055.5% / 34.1% / 57.8% / not reportedRun by Cursor. Sonnet 5.5's scores came from a table Cursor sent to Anthropic
GDPval-AA v2.11,844 / 1,449 / 1,846 / 1,487Elo. Run by Artificial Analysis on a pre-release deployment with a structured-output bug
AA-Briefcase v1.11,811 / 1,359 / 1,822 / 1,483Elo. Same pre-release deployment and bug
Humanity's Last Exam64.5% / 54.9% / 67.7% / not reportedWith tools. Without tools: 56.9%, 43.1%, 64.4%
OSWorld 2.180.1% / 57.0% / 81.8% / not reportedPartial credit. Strict pass rates: 43.5%, 25.6%, 48.7%
Chartography61.6% / 15.6% / 64.4% / 53.6%No tools. With tools Sonnet 5.5 scores 90.2% against Opus's 89.0%
Bar chart of Elo ratings on an axis from 1,000 to 2,000. GDPval-AA v2.1: Sonnet 5.5 1,844, Opus 5.5 1,846, Sonnet 5 1,449, GPT-6 Sol 1,487. AA-Briefcase v1.1: 1,811, 1,822, 1,359 and 1,483.
The two knowledge-work rows are Elo ratings from Artificial Analysis. The axis starts at 1,000 and says so.

Three patterns hold across the table. Sonnet 5.5 beats Sonnet 5 on every row, often by a very wide margin: 395 Elo points on GDPval-AA, 46 percentage points on Chartography, 23 on OSWorld. It trails Opus 5.5 on eight of nine rows, by between 1.7 and 8.2 points where the rows are percentages. And on Terminal-Bench and Chartography, Sonnet 5 scores so low (10.3 and 15.6 per cent) that the size of the jump says as much about the old model as the new one. The page does not explain why Sonnet 5 does so badly on Terminal-Bench 4.0, a version that the system card says was rebuilt to reduce sensitivity to timeouts and harnesses.

The system card's own summary table adds rows the launch page leaves out: SWE-Bench Pro 81.3 per cent against Opus's 89.9, SWE-Bench Multilingual 90.3 against 93.9, SWE-Bench Multimodal 54.3 against 61.4. Two go the other way: HealthBench Professional, 69.2 against Opus's 65.6, and AutomationBench, 44.7 against 42.5.

The cost claims check out, and they are claims against Sonnet 5

The most useful part of the launch page is the four accuracy against cost charts. Each plots every model at five effort levels, from Low to Max, against the average cost per task. The data labels on those charts carry the exact numbers, and every cost claim on the page can be checked against them. We did.

The launch page's cost claims tested against its own chart data. Our arithmetic

Claim on the pageChart dataOur check
Terminal-Bench: Medium beats Sonnet 5's best for less than a tenth of the cost28.8% at $0.83 against 10.3% at $11.62Holds. 7.1 per cent of the cost
FrontierCode: High matches GPT-6 Sol's best for about a fifth of the cost49.4% at $0.42 against 49.3% at $2.07Holds. 20.3 per cent
FrontierCode: 10 points above Sonnet 5 at High, about one fifteenth of the cost49.4% at $0.42 against 39.4% at $6.10Holds. 10.0 points, 6.9 per cent, nearer one fourteenth
CursorBench: Low beats Sonnet 5's best for less than a tenth35.8% at $0.50 against 34.1% at $7.17Holds. 7.0 per cent
AA-Briefcase: Medium beats Sonnet 5's best for about one ninth1,461 at $1.64 against 1,359 at $14.43Holds. 11.4 per cent, nearer one ninth than one eighth

Every comparison in that table is against Sonnet 5 or GPT-6 Sol. None is against Opus 5.5 at matched spend, so we ran that comparison from the same data, picking the nearest-cost points.

Sonnet 5.5 against Opus 5.5 at the nearest matching cost per task. Our selection from the launch page chart data

BenchmarkSonnet 5.5Opus 5.5 at similar cost
Terminal-Bench 4.061.5% at Xhigh, $5.3064.2% at High, $3.88: Opus ahead, for less
CursorBench 4.053.1% at Xhigh, $3.8856.0% at High, $3.97: Opus ahead
FrontierCode 1.149.4% at High, $0.4247.3% at Low, $0.40: Sonnet ahead
AA-Briefcase v1.11,634 at High, $3.951,642 at Medium, $4.40: level

The pattern matches what Anthropic says in prose. Sonnet 5.5 is most distinctive at the cheap end, where Opus has no point to compare. At the expensive end the two share roughly one cost curve, and on Terminal-Bench and AA-Briefcase Sonnet's Max run is the most expensive point on the chart, at $12.54 and $29.19 a task. On AA-Briefcase that is 39 per cent more than Opus at Max ($21.05) for an Elo 11 points lower.

The friendly-name fallacy applies here. "Sonnet" reads as the economy option, and per token it is. Per task, at the settings the headline table uses, it can be the dearer of the two. Effort, not the model name, sets the bill.

Which comparison models were chosen, and who ran the tests

The launch table has one outside comparison, GPT-6 Sol, and prints it on four rows of nine. Where GPT-6 Sol has no public score, the cost charts substitute GPT-5.6 Sol, the previous generation, on Terminal-Bench and CursorBench. The page says so in a note under each chart, which is honest, but a reader glancing at the charts sees an older OpenAI model plotted beside the new Claude one.

The system card names the model that is missing. OpenAI's GPT-6 Astra scores 57.9 per cent on Terminal-Bench 4.0, below Sonnet 5.5, but 64.6 per cent on Terminal-Bench-Science against Sonnet's 59.9, and 65.5 per cent on FrontierSWE v2 against Sonnet's 61.9 and Opus's 62.3. Those rows appear in the system card and not on the launch page. Our Opus 5.5 briefing found the same thing a week ago: Opus lost two benchmarks to GPT-6 Astra. No Google or xAI model appears anywhere on the launch page.

Who produced each launch-table number. Source: system card sections 1.6 and 8

BenchmarkRun byWhat that does not establish
Terminal-Bench 4.0Anthropic, in Claude Code bare modeIndependent replication. The public leaderboard is cited only for other models
FrontierCode 1.1Cognition, the benchmark's authorCosts for Sonnet 5.5: Cognition's files carry none
CursorBench 4.0Cursor, in its production harnessPublic listing: Sonnet 5.5's scores came from a table Cursor sent Anthropic, and its costs are Anthropic's estimates
GDPval-AA, AA-BriefcaseArtificial AnalysisA clean run: the deployment carried a bug since fixed
Humanity's Last Exam, OSWorld, ChartographyAnthropicIndependent replication. OSWorld uses a Claude model as grader on tasks that need one

None of this is unusual. Vendors run most of their own evaluations, and naming the third parties who ran the rest is better practice than many launches manage. It does mean that almost every number in the launch table is one the vendor either produced or selected. The strongest claims, where an outside party ran the test in its own harness, are FrontierCode and CursorBench, and on both Sonnet 5.5 trails Opus 5.5.

The customer quotes and the worked examples

The page carries thirteen customer quotes. Several contain numbers, and they are the most concrete claims about real workloads anywhere on the page. None comes with a method.

Numerical claims in the customer quotes, as printed. Our arithmetic where shown

CompanyClaimNot published
Balyasny Asset Management2,441 finance tasks; about 121k tokens per answer against Sonnet 5's 497k, 76 per cent fewerThe quality scores behind "scored ahead"
Base44118 app builds, level with Opus 5; 3.6 iterations per build against 7.7, 53 per cent fewerHow builds were scored
BoxMore accurate, 2.4 times faster, 12 per cent fewer total tokensTask set and accuracy measure
ZendeskTickets processed 20 per cent faster than current production modelsWhich models, which tickets
SlackBetter on almost all offline evals, about 14 per cent fewer output tokensThe evals
LovableA third fewer tool calls, roughly half the shell runsTask count
Unity90 per cent of tasks completed on its Unity Editor and coding benchmarkThe benchmark or the comparison models
AtlassianRovo Agents up to 30 per cent faster than on Sonnet 5Median, as opposed to best case

These are testimonials from companies that had early access and chose to be quoted, collected by the vendor. That is a method of marketing, not of measurement, and it says nothing about the good faith of any company named. A 76 per cent cut in tokens per answer at Balyasny is a striking figure and may well be true. Nobody outside Balyasny can check it.

The worked examples are the same kind of evidence. Three side-by-side animations built by Sonnet 5 and Sonnet 5.5 from one-line prompts (400 starlings in a murmuration, wind shaping sand dunes, a clock made of 24 clocks) show speed, with no token counts or timings printed. The one internal test described, a 10-slide operating review built from a public company's earnings materials, was judged ready to send by two experts. That is a single task with two judges.

For security teams: more capable, and more refusals

The system card is direct about what changed in offensive capability. Sonnet 5.5 "is not as cyber-capable" as Opus 5.5 or Claude Mythos 5.1, "but it is a significant step up" from Sonnet 5. All the capability figures below were measured with cyber safeguards switched off, which the card says reflects what verified users will get through its Cyber Verification Program.

Cyber capability evaluations, safeguards off. Source: system card section 3.2

EvaluationSonnet 5.5Comparison
Binary Exploitation Benchmark: control-flow hijacks from 831 entry points in 228 OSS-Fuzz projects50Sonnet 5: 3. Mythos 5.1: 81. Opus 5.5: 106
CyScenarioBench, Irregular's multi-stage operations, 10 challenges46.1% solvedSonnet 5: 0.7%. Mythos 5.1: 61.7%. Opus 5.5: 67.6%
ExploitBench, arbitrary code execution178 of 410 runs, 43.4%Cap% of 80 in the nudged arm

Going from 3 control-flow hijacks to 50 is roughly a seventeen-fold increase. That is why Sonnet is now deployed with the three-stage cyber safeguards Anthropic uses on Opus: a probe reading the model's internal activations, a lightweight classifier running on the model itself, and a separate trained classifier model that decides with the probe whether to block. The policy is the same as Opus's. Finding vulnerabilities in source code is allowed. Finding them in compiled binaries is blocked.

The practical consequence is written into the system card in plain terms: "users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks." When a request is blocked, Sonnet 5.5 falls back to Sonnet 5. That happens automatically in Anthropic's own apps. On the API, the developer has to opt in to fallbacks, and without opting in a blocked request is simply blocked. Traffic through other platforms "may experience different behavior".

Three other safeguards arrive with this model. Requests related to developing frontier models, such as kernel work on certain ML accelerators, are blocked and fall back to Sonnet 5. New distillation classifiers block attempts to extract hidden reasoning, with no fallback. And thinking is now tied to the account that produced it, so switching accounts in the middle of a Claude Code session changes behaviour. Biology safeguards are unchanged from Sonnet 5.

The fallback is where the prompt injections landed

Anthropic calls Sonnet 5.5 its "most robust Sonnet-class model to date" against prompt injection, and the system card supports that. The detail inside one of its tables is the most useful finding in the release for anyone deploying a coding agent.

Gray Swan's Shade tool attacked the model in 40 coding scenarios, 200 attempts each. Overall, 3.01 per cent of attempts succeeded against Sonnet 5.5 as deployed, against 19.47 per cent for Sonnet 5. But many of the injected instructions ask for destructive actions such as wiping a disk, which trips the cyber classifier and re-routes the request to Sonnet 5. That happened to 25 per cent of requests, and 12.01 per cent of those re-routed requests were compromised. Sonnet 5.5 itself was compromised in 4 of the 5,901 requests it answered, 0.07 per cent.

Bar chart of prompt-injection attack success in coding on a 0 to 20 per cent scale. Sonnet 5 on its own: 19.47 per cent. Sonnet 5.5 as deployed: 3.01 per cent. Split by the model that answered: requests re-routed to Sonnet 5 after a cyber block, about a quarter of all requests, 12.01 per cent; requests Sonnet 5.5 answered itself, 4 of 5,901, 0.07 per cent.
Gray Swan Shade attacks in coding environments, without prompt-injection probes, from system card section 5.2.2.1. Bar colour shows the model that answered.

Weighting the two parts reproduces the headline: a quarter at 12.01 per cent plus three quarters at 0.07 per cent is 3.05 per cent, against the published 3.01. On our arithmetic, about 1,970 requests went to the fallback and roughly 236 of them were compromised, against 4 compromises of the new model. That is about 98 per cent of the successful attacks. These are derived figures, not ones Anthropic prints, and they assume the share of requests re-routed is close to exactly a quarter.

The finding has limits, and they matter. It is one benchmark in one setting. On Gray Swan's separate indirect-injection benchmark, 11 per cent of Sonnet 5.5's rollouts fell back and the fallback did not make things worse: 1 success in 777 fallback rollouts against 18 in 6,347 others. In computer use, under 1 per cent of requests fell back and neither successful attack came from the fallback. In Anthropic's browser test run through Claude Cowork, no attack succeeded in 110 scenarios, even without safeguards, a first for any model it has tested. The Shade attacker was also trained against earlier Claude models before being run unchanged, and it is our inference, not a finding, that attacks tuned that way may transfer best to an older model.

The mechanism is still worth understanding. A safety control that re-routes a suspicious request to a weaker model can hand an attacker the weaker model, and an injection that looks destructive is exactly the kind that trips the control. Anthropic reported the split rather than only the flattering total, and without it nobody outside the company could have seen this.

Two refusal measures moved the other way. On 61 malicious Claude Code prompts run ten times each, Sonnet 5.5 refused 85.2 per cent against Sonnet 5's 87.9. On 112 malicious computer-use tasks it refused 79.46 per cent against 84.68, the same rate as Opus 5.5. Both figures are without production safeguards, which add the classifiers described above.

What the alignment section says

Anthropic describes Sonnet 5.5 as not advancing the frontier of its models' capabilities, and its Responsible Scaling Policy evaluations found it crosses no new threshold. The alignment work therefore targeted risks that apply at any capability level. On an automated audit of roughly 1,850 scenarios it matches or improves on Sonnet 5 on most measures, and it has the lowest rate of attempted sandbox escapes of any model tested.

Two cautions sit in the same pages. The system card says Sonnet 5.5's "thinking is more illegible than many previous models", which matters because readable reasoning is one of the ways misbehaviour gets caught. And the launch page repeats a line worth quoting in any risk register: Sonnet 5.5 "may have tendencies we haven't found". The system card is also shorter than earlier ones by design: Anthropic says it will condense cards for non-frontier models and omit evaluations that need a lot of human time when they are not critical.

What to check before switching

Take this with you

In the order worth doing

  • Decide the effort level first and test at it. The launch table's scores are mostly at Max, above both the Medium default in Claude Code and the High default on the Claude Platform.
  • Price the switch per task on your own workload. The per-token price is unchanged from Sonnet 5, and cache reads cost the same as on Opus 5.5.
  • If your team does security work, test your real prompts before migrating. The system card warns of more refusals even on benign cybersecurity tasks.
  • On the API, decide whether to opt in to fallbacks. Without it a blocked request fails; with it the request is answered by Sonnet 5.
  • Log the model that actually served each response, not only the one requested, so fallback traffic is visible.
  • Treat destructive tool calls from any coding agent as needing approval. The injection results show the fallback path is where attacks succeeded.
  • If you run Sonnet with thinking off, move to the between_tools setting before switching, as the migration guide requires.
  • Check whether switching accounts mid-session is part of any workflow, because preserved thinking is now tied to the account.
  • Keep Opus 5.5 for open-ended work. Anthropic's own text says it remains clearly stronger there.

The question this launch raises

Sonnet 5.5 is a well-documented release. The benchmarks are all printed, the conditions are in the footnotes, the cost data sits behind every chart, and the one safety number that is unflattering to Anthropic's own design is broken out in its system card rather than buried in a total. Most launches give far less.

The question it leaves is for the buyer, not the vendor. If your agents run at the default effort, which model and which setting are your numbers actually coming from, and when a classifier steps in, which model answers?

Sources

  1. PrimaryIntroducing Claude Sonnet 5.5, 28 September 2026: benchmark table, the four accuracy against cost charts (read from their data labels), pricing, customer quotes, worked examples, safety and availability. Read in fullAnthropicaccessed 2026-09-28
  2. PrimarySystem Card: Claude Sonnet 5.5, 148 pages: safeguards, cyber evaluations, agentic safety and prompt injection, alignment and capability method (sections 1, 3, 5, 6 and 8 read in full)Anthropicaccessed 2026-09-28
  3. PrimaryModels overview: context window, maximum output, knowledge cutoff, default effort, retirement date and platform model IDsAnthropicaccessed 2026-09-28
  4. PrimaryClaude Sonnet product page: availability on Claude.ai, US-only inference at 1.1 times pricing, caching and batch discountsAnthropicaccessed 2026-09-28
  5. Reported byOur briefing on Claude Opus 5.5, 23 September 2026, for the Opus comparison and the earlier safeguard routingpk-sharma.comaccessed 2026-09-28

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.