ZYVOPMulti-Platform Sync
SeriesAI NewsWhy ZyVOPJoin Discord
LoginGet Started
ZYVOP
The Developer Publishing Hub
PrivacyTermsGuidelinesDMCACommunity
© 2026 ZyVOP
HomeClaude Opus 5.5 Just Landed

Claude Opus 5.5 Just Landed

Claude Opus 5.5 matches Fable 5.1 on most tasks, cuts run costs 40%, prices tokens lower, and ships with expanded biology and cybersecurity safeguards.

Sanju Singh
Sanju Singh
September 22, 2026•
5 min read
Claude Opus 5.5 Just Landed
#Anthropic#llm-benchmarks#Claude Opus 5.5#AI pricing#AI models

Anthropic released Claude Opus 5.5 on September 22, 2026, its first model since CEO Dario Amodei published "We Must Pace the Frontier," an essay from about a week earlier arguing that safety practices need to keep pace with capability jumps. The model was evaluated before launch by outside groups including METR and Frontier Design.

The pitch

Opus 5.5 performs roughly at the level of Anthropic's top public model, Claude Fable 5.1, on most tasks, while costing 40% less to run than the outgoing Opus 5. Anthropic is candid that at this capability tier, benchmark gaps are becoming less reliable as a stand-in for real-world differences.

Benchmarks

Verified against Anthropic's own comparison table.

Benchmark

Opus 5.5

Fable 5.1

Opus 5

GPT-6 Astra

GPT-5.6 Sol

Terminal-Bench 4.0

66.4%

55.8%

52.3%

57.9%

37.3%

FrontierCode v1.1

54.4%

50.3%

48.0%

53.3%

47.5%

CursorBench 4.0

57.8%

51.8%

46.6%

-

41.7%

GDPval-AA v2.1 (Elo)

1846

1735

1708

1542

1588

Humanity's Last Exam

67.7%

65.6%

63.6%

57.2%

-

AutomationBench

40.0%

31.4%

26.9%

41.4%

28.8%

Terminal-Bench-Science 0.1

58.7%

52.6%

29.0%

64.6%

22.4%

OSWorld 2.0 (computer use)

81.8%

80.7%

74.0%

-

-

Chartography (chart reading)

89.0%

88.4%

83.4%

-

-

These numbers match Anthropic's published figures exactly, though OSWorld 2.0 and Chartography are scored under partial-credit and tool-assisted conditions per Anthropic's own footnotes, and GPT-6 Astra and GPT-5.6 Sol are absent from both rows, so those two comparisons are Claude-only rather than a full field test.

The two bolded losses to GPT-6 Astra have different causes. On Terminal-Bench-Science 0.1, Anthropic's general safeguard footnote applies: when Opus 5.5's safety systems intervened during testing, biology and frontier-AI-research tasks were completed by the older Opus 5 instead, which likely dragged the score down.

On AutomationBench, safeguards are still the cause, just structured differently. Those results were run by Zapier without fallback models, so any safeguard intervention counted as an automatic failure rather than being completed by a substitute model, as on Terminal-Bench-Science, producing a lower score than Opus 5.5 would achieve in practice.

Coding: the headline act

One tester completed a 680,000-line code migration in under a day. Asked to cut load times across a web app, Opus 5.5 succeeded 39 of 40 times without breaking behavior, where Opus 5 made smaller improvements that also altered the app's behavior. A 200,000-line codebase audit took Opus 5.5 under three hours versus Opus 5's 20-plus hours, at roughly 40% of the token cost.

Translating HAProxy from C to Rust, Opus 5.5 passed nearly all regression tests in 9.5 hours against Fable 5.1's 12, at about half the cost. Against GPT-6 Astra, it matches Terminal-Bench 4.0 for around 40% of the price and beats FrontierCode for about 20% of the price. In a separate test where several Claude models built a game from a single prompt, Opus 5.5's build scored highest on graphics and polish.

Customers quoted in the release, including GitHub, Box, and Optiver, independently described similar drops in tokens per task. Stripe is also quoted in the release, though on session efficiency completing a 40-pull-request rebase rather than on token counts.

Security. Opus 5.5 matches or beats Opus 5 against prompt-injection attacks across coding, tool use, computer use, and web browsing, and ties Fable 5.1 for the lowest prompt-injection success rate in third-party testing by the AI security firm Gray Swan. For long, unattended coding sessions, Anthropic also pairs the model with a pre-action classifier, an auditable open-source sandbox, and code review meant to catch vulnerabilities before they merge.

Not just code

In a blind earnings-report test where any invented figure counted as an automatic fail, 16 of 18 Opus 5.5 attempts passed; neither Fable 5.1 nor Opus 5 passed on any attempt. Walleye Capital, an early enterprise tester, said Opus 5.5 caught an error in the firm's own evaluation instructions that no prior model had noticed, a detail confirmed directly in Anthropic's writeup.

In a separate merger-analysis test, Opus 5.5 and Opus 5 each built an Excel financial model and turned it into an executive presentation. Both reached the same conclusion, but Opus 5.5 finished in 63 minutes versus 93 for Opus 5, at half the cost and with a cleaner model.

Two more data points came from named enterprise testers. At its lowest effort setting, Deloitte says Opus 5.5 caught 72% of known bugs in code review, against 56% for Opus 5 at high effort, and Hebbia's finance-workflow rubric put Opus 5.5 at 86.6% coverage versus 60.3% for Opus 5, alongside the model's best-ever citation recall.

Writing style overhaul

Anthropic says Opus 5.5 leads with the important information first and uses less jargon, addressing the most common complaint about Opus 5. Box reports roughly a two-thirds cut in token usage on content-heavy workloads as a result.

Safety

On Anthropic's roughly 2,000-scenario behavioral audit, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior and is the company's strongest model to date on most measures of honesty. In a new evaluation built to test how often a model tries to cross containment boundaries, Opus 5.5 attempted this about 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it did make was low severity and self-reported.

Because Opus 5.5 is now comparable to Mythos 5.1 in biology and cybersecurity, it ships under safeguards similar to Fable 5.1's: most flagged cybersecurity tasks are re-routed to the older Opus 4.8, and flagged biology tasks go to Opus 5, with full biology capability available to vetted organizations through Anthropic's Life Sciences Verification Program.

Anthropic says this reflects real capability rather than caution alone: in biology, Opus 5.5 improved on a long-horizon molecular-design evaluation run with Dyno Therapeutics, and outside red-teamers rated its scientific novelty on par with the best model they had tested.

On the cybersecurity side, Anthropic is expanding its Cyber Verification Program to three tiers of increasingly permissive access for vetted practitioners, up to and including the restricted Mythos models. Anthropic also reports signs that Opus 5.5 often suspects it is being evaluated, which complicates its ability to predict how the model will behave across the real-world settings it gets deployed into. Evaluation awareness is a known open problem across the AI industry, not unique to Anthropic.

Opus 5.5 also launches with preserved thinking, the anti-distillation safeguard introduced with Fable 5.1, which blocks API users from editing Claude's prior context to extract its reasoning. It applies to API accounts created on or after August 31, 2026, and extended thinking can no longer be fully disabled.

Pricing & availability

Opus 5.5

Opus 5

Input (per 1M)

$4

$5

Output (per 1M)

$20

$25

Cache reads (per 1M)

$0.20

$0.50

Cache writes (per 1M)

$5

$6.25

At standard settings, Opus 5.5 also generates output more than 30% faster than Opus 5. A separate fast mode in Claude Code and the Claude Platform trades some of that headroom for up to 2.5x the speed, priced at $8 per million input tokens and $40 per million output tokens.

Live now on the Claude Platform (claude-opus-5-5), AWS, Google Cloud, and Microsoft Azure. Pro/Max/Team subscribers and seat-based Enterprise plans get bumped five-hour limits plus a bankable rate-limit reset. Sonnet 5.5 and Haiku 5.5 are coming in the following weeks.


A note on accuracy

The "smaller improvements that also altered the app's behavior" line about Opus 5 comes straight from Anthropic's own comparison, not an independent claim, since Anthropic is grading its own predecessor here. Worth keeping that context in mind when reading vendor-published benchmarks.

The same goes for every customer statistic in this piece, Deloitte, Hebbia, GitHub, and the rest: those were selected and published by Anthropic for the launch, not gathered independently, so treat them as testimonials rather than a representative sample.

Comments (0)

Join the discussion by logging into your account.

Sanju Singh
Sanju Singh

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Sanju Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.

Sanju Singh
Like
Love
Clap
Fire
Party
Wow

More from Sanju Singh

View profile

Should You Still Learn to Code Now That AI Can Write It?

Jensen Huang says AI ended the need to learn to code. But Anthropic's randomized trial, METR's productivity study, and Stanford's labor data point somewhere else, toward who really benefits from AI and who just thinks they do.

4 minSep 25

Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation

Alibaba open-sourced Qwen-Image-2.1, a 7B-parameter model unifying generation and editing with native transparency and up to 10 reference images. Day-zero framework support is strong, but it ships under a non-commercial license with no independent benchmarks yet.

5 minSep 21

Cloudflare Quick Tunnels: One Command, Three Hard Limits

Quick Tunnels expose localhost in one command, no signup required. But they cap at 200 concurrent requests, drop Server-Sent Events, and carry no SLA. Here's the mechanics, a Node helper that reads the tunnel URL properly, and when to stop using them.

13 minSep 19

From 64MB to 16GB: How Software Got So Hungry

Microsoft's published minimum RAM requirement rose roughly 256x between 2001 and 2024. Here's the paper trail behind that number, a correction to the most-repeated Tauri benchmark, and a way to measure your own Electron app's memory footprint tonight.

6 minSep 18

Neural Networks, Explained Simply - Part 2: How Neural Networks Actually Learn

Part 2 of our Neural Network Series: how a neural network starts out guessing randomly and learns from its mistakes through training and backpropagation. A plain-language look at how the correction cycle actually works, no calculus needed.

3 minSep 17