
Anthropic released Claude Haiku 5.5 on October 7, 2026, and its headline claim is that the model costs around 75% less to run than Haiku 4.5 on average. That number is real but conditional, and the model is not a drop-in swap. Several request shapes that worked on Haiku 4.5 now return a 400 error. This post covers the price math, the benchmark picture, the migration changes, and a routing plan you can ship this week.
The short version
Item | Claude Haiku 5.5 |
|---|---|
Model ID |
|
Released | October 7, 2026 |
Context window | 1M tokens, up from 200K on Haiku 4.5 |
Max output | 128K tokens, up from 64K on Haiku 4.5 |
Price per 1M tokens | $0.10 input and $0.50 output for prompts up to 100K tokens, $0.50 and $2.50 above that |
Thinking | Adaptive, on by default, default effort |
Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
Sources: the model overview, the what's new page, the migration guide, and the launch post.
Where the 75% comes from
The launch post lists the price change line by line. Every row drops by the same ratio: 90% for prompts up to 100,000 tokens and 50% above that.
Price per 1M tokens | Haiku 4.5 | Haiku 5.5, up to 100K | Haiku 5.5, over 100K |
|---|---|---|---|
Input | $1.00 | $0.10 | $0.50 |
Output | $5.00 | $0.50 | $2.50 |
Cache reads | $0.10 | $0.01 | $0.05 |
Cache writes | $1.25 | $0.125 | $0.625 |
Anthropic says that about 90% of Haiku 4.5 requests fell into the cheaper bucket, and that the 75% average also accounts for a new tokenizer that uses slightly more tokens per task. The docs put a number on that: the same text counts as approximately 30% more tokens, with the exact increase depending on content.
Request share is not spend share. Long prompts cost more per request, so they carry a bigger slice of a typical bill than 10%. Because every price row drops by the same ratio, your savings depend on one variable: the fraction f of your Haiku 4.5 spend that lands in requests over the 100K threshold. The table below is our own arithmetic from Anthropic's published prices and the docs' 30% figure. It is not an Anthropic number.
Long-prompt share of old spend (f) | New cost as a fraction of old | Savings |
|---|---|---|
0% | 0.130 | 87% |
10% | 0.182 | 82% |
25% | 0.260 | 74% |
50% | 0.390 | 61% |
100% | 0.650 | 35% |
Working backward, a 75% average is consistent with roughly 23% of old spend sitting in long prompts, if the token inflation applies evenly. That is our inference, not a disclosed figure. The practical point stands: compute your own f instead of assuming 75%. If you log token counts per call, one query gets you an estimate.
-- Assumes llm_calls(model text, input_tokens int, output_tokens int).
-- Token counts are from Haiku 4.5, so scale by 1.3 to guess which calls
-- cross 100K on Haiku 5.5. Ignores cache pricing, so treat it as an estimate.
SELECT
SUM(CASE WHEN input_tokens * 1.3 > 100000
THEN input_tokens * 1.00 + output_tokens * 5.00 ELSE 0 END)
/ NULLIF(SUM(input_tokens * 1.00 + output_tokens * 5.00), 0)
AS long_prompt_spend_share
FROM llm_calls
WHERE model LIKE 'claude-haiku-4-5%';One trap deserves its own warning. Under the 30% figure, a prompt that measured about 77,000 tokens on Haiku 4.5 comes out near 100,000 on Haiku 5.5, which is where the higher price tier starts. The migration guide tells you to count prompts with the model set to claude-haiku-5-5 rather than reuse old counts, and that is the only reliable way to know which side of the line a prompt lands on.
Here is a worked example. A job sends 50 million input tokens and 5 million output tokens a month on Haiku 4.5, which costs $75. On Haiku 5.5 the same text counts as about 65 million and 6.5 million tokens. If every prompt stays under the threshold, the bill is $9.75, an 87% drop. If every prompt is over it, the bill is $48.75, a 35% drop. The calculator below produced both numbers, and we ran it on Node 22.
// cost.ts
const THRESHOLD = 100_000;
const INFLATION = 1.3; // docs: approximately 30% more tokens, content dependent
export type Call = { inputTokens: number; outputTokens: number }; // Haiku 4.5 counts
export function haiku45Cost(c: Call): number {
return (c.inputTokens * 1.0 + c.outputTokens * 5.0) / 1e6;
}
export function haiku55Cost(c: Call): number {
const input = c.inputTokens * INFLATION;
const output = c.outputTokens * INFLATION;
const [pin, pout] = input > THRESHOLD ? [0.5, 2.5] : [0.1, 0.5];
return (input * pin + output * pout) / 1e6;
}
export function savings(calls: Call[]): number {
const before = calls.reduce((s, c) => s + haiku45Cost(c), 0);
const after = calls.reduce((s, c) => s + haiku55Cost(c), 0);
return 1 - after / before;
}Output tokens can rise for a second reason. Haiku 5.5 has adaptive thinking on by default with a default effort of medium. If your Haiku 4.5 calls ran without thinking, the migration guide says to pick a lower effort level. So benchmark cost per completed task at low and medium effort, not cost per token. Batch processing still takes 50% off input and output, and cache reads are down to $0.01 per million tokens in the short bucket.
What the benchmarks say
Anthropic reports the scores below in its launch post. They are vendor-run. The methodology lives in the system card, which we did not audit, and we did not reproduce any of these runs.
Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
GDPval-AA v2.1 (knowledge work) | 1620 | 735 | 1437 | 1840 |
AA-Briefcase v1.1 (knowledge work) | 1578 | 614 | 1336 | 1824 |
OSWorld 2.1, offline subset (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
Humanity's Last Exam, no tools | 45.9% | 10.2% | n/a | 56.9% |
Humanity's Last Exam, with tools | 57.4% | 18.7% | n/a | 64.5% |
Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
FrontierCode 1.1 Main (agentic coding) | 46.4% | n/a | 42.4% | 52.1% at Xhigh effort |
Chartography, no tools (visual reasoning) | 46.4% | 6.4% | 29.1% | 61.6% |
The jump over Haiku 4.5 is large. On OSWorld 2.1 the score moves from 15.7% to 72.4%, and on GDPval-AA v2.1 from 735 to 1,620. The 0.0% for Haiku 4.5 on Terminal-Bench 4.0 is a floor effect: it says the older model could not do the task, not how big the gap is.
The gap to Sonnet 5.5 is just as informative. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% against 70.6%. Anthropic says as much itself: it points to Sonnet 5.5 and Opus 5.5 for complex agentic coding and describes Haiku 5.5 as best suited to narrowly scoped work such as compaction, summarization, and subagents.
Outside Anthropic's own table, The New Stack reports that Artificial Analysis scores Z.ai's GLM-5.3-Flash at 1,647 on GDPval-AA v2.1, slightly above Haiku 5.5's 1,620. It also reports that Alibaba's Qwen3.7 Flash lists at $0.03 per million input tokens and $0.13 per million output tokens on its international service, for inputs up to 32,000 tokens. Haiku 5.5 is the cheapest Claude, not the cheapest small model on the market.
Customer reports in the launch post point the same direction, with caveats. Asana measured over 30% lower latency and up to 2.5x faster inference per agent turn, but against whatever model it runs today, not necessarily Haiku 4.5. Box saw a score 11 points higher than Haiku 4.5 at about half the latency. AlphaSense measured 0.84 versus 0.76 on 400 queries. HubSpot reported 92.8% on its CRM simulation suite.
Speed needs a precise reading. The launch post calls Haiku 5.5 Anthropic's fastest model at each model's standard speed, but notes that it runs less quickly than the Opus models in Fast Mode. We found no tokens-per-second figure in the announcement or the model docs, so measure latency on your own prompts before you promise anyone a number.
Where it fits: a routing plan
Anthropic positions Haiku 5.5 for summaries, compactions, database queries, and classification, as a subagent next to Sonnet 5.5 or Opus 5.5 on coding work, and for speed-sensitive jobs like live customer support and browser use. The model docs summarize it as classification, routing, extraction, and subagent tasks. A router built on those lines looks like this.
flowchart LR
A[Incoming task] --> B{Haiku 5.5 triage at low effort}
B -->|classify, extract, summarize| C[Haiku 5.5 answers directly]
B -->|narrow parallel work| D[Haiku 5.5 subagents]
B -->|complex multi-step coding| E[Sonnet 5.5 or Opus 5.5]
D --> F[Lead model merges results]
E --> F
C --> G[Response]
F --> GThe pattern has early support. Cognition says Haiku 5.5 as a sidekick, with Opus 5.5 as lead, holds a FrontierCode score of 66.2 in Devin Fusion while cutting cost and latency. That is a customer-reported result on one product, so read it as a signal to test the pattern, not as a guarantee.
One design consequence follows from the pricing. Anthropic notes that Haiku 5.5 is especially good value for prompts up to 100,000 tokens. That makes a compaction step cheap and attractive: let Haiku 5.5 compress long context before a larger model sees it, and you stay in the low price tier on the cheap model while shrinking the bill on the expensive one.
For agents that drive a screen, Haiku 5.5 supports the browser use tool (browser_toolset_20260801) on the Claude API and Google Cloud, and Anthropic is adding beta support for computer use and browser use to its Python and TypeScript SDKs.
Migration: the changes that return a 400
The migration guide lists each change. This table condenses it.
Haiku 4.5 behavior | On Haiku 5.5 | Fix |
|---|---|---|
Model ID | Different ID | Use |
| 400 error | Use |
| Non-default values return 400 | Omit all three |
Assistant prefill as the last message | 400, even with thinking off | End |
Computer use via | 400 on Claude API and Google Cloud | Use |
First content block is the answer | Responses can start with thinking blocks | Select blocks by |
Editing | 400 if you send thinking blocks back | Keep conversations append-only |
Priority Tier commitment | Not supported | Plan capacity separately |
Safety classifiers | Can return | Handle refusals in your client |
Two defaults also changed. Thinking blocks now come back with an empty thinking field and only a signature, and you set display to summarized if you want the summary. A forced tool_choice returns the tool call with no thinking block, so use auto and say in the prompt when to call the tool if you want the model to think first.
Thinking blocks are also account-bound. They work only in the account that produced them or a linked account, so a service that replays stored conversations through a different customer account loses that reasoning without an error. Thinking tokens count toward max_tokens, so a small limit can end the response after a thinking block and before any text.
The append-only rule has an exception for older accounts. On accounts created before August 31, 2026, the 400 appears only on requests that set thinking.block_binding.prefix_mismatch_behavior, according to the migration guide. Accounts created after that date are not covered by the exception, so assume the rule applies to you.
Here is a minimal call that follows the request shape in the guide. It is a sketch for a NestJS service or any Node 22 runtime, and we have not run it against the live API, so test it with your own key first.
type Effort = "low" | "medium" | "high";
export async function askHaiku(prompt: string, effort: Effort = "low") {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY!,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({
model: "claude-haiku-5-5",
max_tokens: 2048, // thinking tokens count toward this
thinking: { type: "adaptive" },
output_config: { effort },
messages: [{ role: "user", content: prompt }], // no prefill, no temperature
}),
});
if (!res.ok) throw new Error(`Anthropic ${res.status}: ${await res.text()}`);
const data = await res.json();
if (data.stop_reason === "refusal") return { kind: "refused" as const };
if (data.stop_reason === "max_tokens") return { kind: "truncated" as const };
const text = data.content
.filter((b: { type: string }) => b.type === "text")
.map((b: { text: string }) => b.text)
.join("");
return { kind: "ok" as const, text, usage: data.usage };
}Safeguards you should know about
Anthropic reports major improvements across almost all of its alignment evaluations relative to Haiku 4.5, including far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse. The details are in the system card.
The cybersecurity safeguards are more restrictive than Haiku 4.5's and somewhat less restrictive than those on other recent models. They permit a wider range of defensive tasks than the Sonnet 5.5 safeguards, but they still block penetration testing. The biology safeguards match Sonnet 5, Sonnet 5.5, and Opus 5. Organizations that need wider access can apply to the Cyber Verification Program or the Life Sciences Verification Program.
If you build security tooling on a small model, this matters more than any benchmark. A classifier-driven refusal in the middle of an agent loop needs a handler, a fallback model, and an alert, because the platform will not fall back for you.
Also announced the same day
Anthropic halved the price of Sonnet 5.5 cache reads, from $0.20 to $0.10 per million tokens, and says that makes Sonnet 5.5 around 20% cheaper on most agentic work. It is also rolling out a monthly API credit for subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. The credits work on any model on the Claude Platform.
A rollout plan
Step | Action | Pass condition |
|---|---|---|
1. Count | Run token counting on a sample of real prompts with | No prompt unexpectedly crosses 100,000 tokens |
2. Fix requests | Apply the migration table on a branch and replay recorded traffic | No 400 errors on replay |
3. Evaluate | Run your own eval at low and medium effort against Haiku 4.5 | Equal or better accuracy at a lower cost per completed task |
4. Shadow | Send a slice of production traffic to both models and compare | Latency and refusal rate inside your limits |
5. Route | Move narrow tasks, keep complex coding on Sonnet 5.5 or Opus 5.5 | Blended cost per task falls without a quality regression |
What we could not verify
Three limits apply. We did not reproduce any benchmark, so every score above is Anthropic's own or a customer's own. The 75% figure is an average over Anthropic's request mix, and under our 30% token assumption yours can land anywhere from about 35% to 87% depending on prompt length, with extra thinking tokens pushing it lower. And when we fetched Anthropic's general Haiku product page on October 8, 2026, it still described Haiku 4.5, so treat the launch post and the platform docs as the record.
Bottom line
Claude Haiku 5.5 is a real step up for high-volume work, and the price cut is large enough to change which jobs are worth sending to a model at all. The honest caveats are that the average savings depend on your prompt lengths, that a bigger tokenizer and default thinking can erode them, and that several request shapes need code changes before the first call succeeds. Count tokens, fix the requests, evaluate at low effort, and route narrow tasks to it first.
Sources
Source | Used for |
|---|---|
Pricing table, 75% claim, benchmarks, customer results, safeguards, Sonnet 5.5 cache price cut, API credits | |
Model ID, release date, context window, output limits, batch discount, platforms | |
Haiku 4.5 limits, tokenizer change, adaptive thinking | |
Breaking changes, request shapes, refusals, Priority Tier | |
Competitor scores and pricing context |
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!