
On Monday, September 28, Anthropic released Claude Sonnet 5.5 at the same price as Sonnet 5. The company says it finishes the same work for up to 30% less, mostly because it uses fewer tokens and makes fewer tool calls.
Anthropic pitches Sonnet as the everyday model, strongest at well-scoped tasks like fixing bugs and producing documents, slides, and spreadsheets. Opus 5.5, released about a week earlier, is the bigger model in the 5.5 family.
The quick facts
Released: September 28, 2026, as the second model in the Claude 5.5 family.
Price: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads, all unchanged from Sonnet 5.
Speed and cost: output is more than 30% faster than Sonnet 5, and cost per task is down by up to 30% in Anthropic's testing.
Where to get it: the Claude apps, the Claude Platform (
claude-sonnet-5-5), AWS, Google Cloud, and Azure, with zero data retention available.
A few terms
Token: a small chunk of text, roughly a word or part of one. Usage is billed by the token.
Effort: a setting for how long the model works before it answers. Higher effort usually costs more.
Tool call: when the model uses another program, like running a command or searching the web.
Benchmark: a standard test used to compare models.
Judge it by cost per task
Since the price per token didn't change, the savings come from using fewer tokens. That's a better way to compare models than the price list. What you pay for is a finished job, including retries and wasted steps, and a model that wanders can cost more than a pricier one that goes straight to the answer.
Customers reported the same effect. Balyasny Asset Management ran 2,441 finance tasks, and Sonnet 5.5 scored ahead of Sonnet 5 while using about 121,000 tokens per answer against about 497,000. Slack reported about 14% fewer output tokens, and Box 12% fewer total tokens.
Sonnet vs. Opus
Price per 1M tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
Input | $2 | $4 |
Output | $10 | $20 |
Cache reads | $0.20 | $0.20 |
Cache writes | $2.50 | $5 |
Opus costs double on everything except cache reads. Cache reads and writes are what you pay to store and reuse text the model has already processed. For a made-up job with 500,000 input tokens and 100,000 output tokens, that's $2 on Sonnet and $4 on Opus. Real jobs won't use identical token counts on both models, so measure your own.
Kevin Ngo, a creative coder at Creator, said he'd feel confident letting Sonnet 5.5 implement a game once Opus 5.5 has set the architecture. That plan-with-Opus, build-with-Sonnet split is one way to use both.
Effort settings
Sonnet 5.5 has five effort levels: Low, Medium, High, Xhigh, and Max. Higher effort means the model works longer, which usually raises both cost and score. Claude Code and the Claude apps default to Medium, and the Claude Platform defaults to High.
Anthropic's main claim is that on several benchmarks, Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. Terminal-Bench 4.0 at Medium and CursorBench 4.0 at Low both did it for under that, and AA-Briefcase at Medium for about one ninth.
More effort doesn't always help, though. On FrontierCode, Sonnet 5.5 scored 52.1% at Xhigh but only 46.2% at Max. Anthropic's footnote says that at Max the model ran Claude Code's code-review skill more often, and in two cases Cognition examined, that caused a timeout or edits outside the task's scope, which the test penalizes.
The scorecard
Benchmark | What it tests | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|---|
Terminal-Bench 4.0 | Multi-step coding work in a command line | 70.6% | 10.3% | 66.4%* |
FrontierCode 1.1 | Whether a code change could be merged | 46.2% (Max), 52.1% (Xhigh) | 42.4% | 54.4% |
CursorBench 4.0 | Tasks from real Cursor coding sessions | 55.5% | 34.1% | 57.8% |
GDPval-AA v2.1 | Real-world work across many jobs | 1844 | 1449 | 1846 |
AA-Briefcase v1.1 | Long knowledge-work tasks | 1811 | 1359 | 1822 |
Humanity's Last Exam | Hard multi-subject questions, with tools | 64.5% | 54.9% | 67.7% |
OSWorld 2.1 | Using a computer, partial credit | 80.1% | 57.0% | 81.8% |
Chartography | Reading charts, no tools | 61.6% | 15.6% | 64.4% |
*Opus 5.5's Terminal-Bench score is at Xhigh effort, which is its best. Only some rows in Anthropic's table say which effort setting was used.
Sonnet 5.5 sits close to Opus 5.5 across the table. The GDPval-AA gap is two points and the CursorBench gap is about two. Opus scores higher on every row except Terminal-Bench.
Two results deserve a second look. The Terminal-Bench jump from 10.3% to 70.6% is large enough that you should check how the test was run before quoting it. And the two knowledge-work scores, run by Artificial Analysis, came from a pre-release deployment with a bug affecting structured outputs. Anthropic expects any effect to be small and to understate the model, and says the bug has been fixed.
Anthropic's table also lists OpenAI's GPT-6 Sol, but OpenAI recently fixed an image-understanding bug that some of those outside scores may not reflect, so this post leaves that comparison out.
Coding
At High effort on FrontierCode (a setting the table doesn't show), Sonnet 5.5 scores 10 points above Sonnet 5 at the same setting, at about one fifteenth of the cost per task. Early testers said it gets its bearings in a codebase quickly and groups tool calls together more than Sonnet 5, which means fewer steps.
Company reports:
Epic Games: in early testing, it managed tens of thousands of lines of gameplay architecture and handled multi-hour tasks with less detailed prompting.
Base44: across 118 real app builds, its apps scored level with Opus 5. It needed 3.6 iterations per build, where Opus 5 needed 7.7.
Lovable: in its own coding tests, about a third fewer tool calls and roughly half the shell runs to finish a task.
These come from customers Anthropic chose to feature, so treat them as encouraging rather than conclusive.
Knowledge work and design
GDPval-AA covers real-world tasks across 44 occupations and nine major industries. Sonnet 5.5 lands nearly level with Opus 5.5 there, and about 400 points above Sonnet 5.
Anthropic says the model writes more clearly than the previous generation, and testers pointed to a knack for design, including slide decks built from a template that need little editing. In one Anthropic test, the model turned a public company's earnings materials, call transcripts, and a slide template into a 10-slide operating review. Two experts judged the first draft ready to send as is.
Zendesk tested it on hundreds of real support cases and saw fewer wrong decisions and tickets processed 20% faster. Box said it rechecks figures against source documents and catches errors Sonnet 5 missed.
Where Opus still wins
On several tests, Sonnet 5.5 at Max effort performs comparably to Opus 5.5. Even so, Anthropic says Opus 5.5 is clearly stronger at complex, open-ended work that needs sustained judgment, and it cautions that benchmarks capture only one side of a model.
Sonnet 5.5 pairs best with Opus at lower effort, where it costs less per task. At higher effort it performs comparably at a similar cost, so running Sonnet at top effort to stand in for Opus won't save much.
The next part is opinion, not Anthropic's claim: use Sonnet for well-scoped, high-volume work where you can check the output, and Opus for ambiguous, high-stakes work where a bad call is expensive.
Safety and safeguards
Anthropic's automated audit covers roughly 1,850 scenarios. Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty, and it comes close to Opus 5.5 on newer sandbox-escape tests, though Opus still does slightly better overall. Anthropic says it found no evidence that Sonnet 5.5 pursues goals that conflict with what the user wants, while acknowledging that no set of tests catches every failure.
Three safeguards change how the model behaves:
Cybersecurity: Anthropic says Sonnet 5.5's cyber skills are a large step up from Sonnet 5's, so it's the first Sonnet to launch with cyber safeguards similar to those on its most capable models. Routine work like fixing bugs in your own code is unaffected, but higher-risk security tasks visibly fall back to Sonnet 5. Defenders will soon be able to apply to an expanded Cyber Verification Program for tiered access to more advanced capabilities.
Biology: the safeguards are the same as Sonnet 5's and target a narrow set of harmful requests, though some microbiology and virology requests may be flagged by mistake. Organizations can apply to the Life Sciences Verification Program.
Distillation: this is when attackers use thousands of fake accounts to copy a model's abilities at scale. Sonnet 5.5 is the first Sonnet with classifiers built to prevent this kind of reasoning extraction, and it expands preserved thinking, which ties Claude's thinking to the account that created it.
Moving over: a checklist
Change the model ID to
claude-sonnet-5-5.If you run Sonnet with thinking off, switch to the new
between_toolssetting before you move. Anthropic has a migration guide.Set your effort level on purpose instead of accepting the default.
If you move conversations between accounts, read the preserved thinking docs first.
Should you switch?
If you're on Sonnet 5, probably: same price, faster output, and fewer tokens per job. But test it first. Pick ten tasks you run every week and run them on both models, then compare total cost per finished task, time to done, and how much cleanup the output needs. Try Low and Medium effort before settling on High, and if you already use Opus, try Sonnet as the builder under an Opus plan.
Anthropic says Haiku 5.5, its small, fast model for high-volume work, will arrive in the coming weeks.
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!