By Mengchen, from QbitAI (Ao Fei Si)
QbitAI | Official Account QbitAI
Claude Haiku 5.5 is released — a big comeback for small models!
In terms of benchmarks alone, it directly surpasses the two former "slaughter lines" DeepSeek V4.1Flash and GLM-5.3-Flash, becoming the new gatekeeper of small models.
Looking at the official benchmarks, well, this was aimed straight at GPT-6 Luna, beating Luna item by item.
Since the pelican riding a bicycle test has basically been saturated, the latest one thrown onto the battlefield is the running zebra test.
Haiku 5.5's price even fully matches GPT-6 Luna.
The previous generation Haiku 4.5 was released exactly one year ago, at 10 times the price of 5.5.
However, Haiku 5.5's price comes with conditions: for requests with prompts within 100,000 tokens (accounting for 90% of Haiku 4.5 requests), input and output prices are cut by 90% directly; for the portion exceeding 100,000 tokens, the cut is only 50%.
This way, except for the cache-hit input item, Haiku 5.5's price is also cheaper than DeepSeek-V4.1 Flash.
Now Opus 5.5 handles complex reasoning, Sonnet 5.5 handles general execution, and Haiku 5.5 handles volume and speed — all the talents have arrived, only Fable is missing.
One can only say that if OpenAI next door doesn't work harder, Claude has no need to squeeze out the big tube of toothpaste that is Fable 5.5.
It feels like returning to the days of Intel and AMD competing over CPUs.
New tokenizer, consumes more tokens
Haiku 5.5 is the first Haiku model to support effort adjustment, inheriting the series' five-tier configuration from low to max.
The previous generation Haiku 4.5 basically turned in a blank sheet on computer operation and coding, while this generation has directly evolved.
Looking at OSWorld 2.1 (testing agents operating real computers to complete multi-step tasks): the Low tier achieves 42.0% accuracy at a per-run cost of $0.07; the Max tier 72.4% at $0.61. Haiku 4.5 only had 15.7% at $1.45 on the Max tier,
That is to say, 5.5's Low tier is more accurate than 4.5's Max tier, and half the price.
GDPval-AA v2.1 (testing real professional work across 44 occupations) shows a similar curve: Low tier Elo 1125 at $0.01; Max tier 1620 at $0.87. 4.5's Max tier only got 735 at $0.24.
However, Haiku 5.5 has switched to a new tokenizer (the same as Sonnet 5.5 and Opus 5.5), so the same task consumes more tokens, meaning the actual money saved is a teeny bit less than the listed discount.
Especially code, tables, andnon-English contentmay inflate even more.
In the AA benchmark, although Haiku 5.5's API price is cheaper, its "average cost per completed task" did not reach the ideal range.
Can't Replace Sonnet
Many people's expectation for Haiku 5.5 is that it could replace Sonnet 5.5 on some tasks, further driving down costs.
But no miracle happened this time — a small cup is still a small cup.
On Terminal-Bench 4.0, Haiku 5.5 scored 39.2% while Sonnet 5.5 scored 70.6%. The gap in complex multi-step coding, cross-file refactoring, and long-horizon autonomous planning is very clear.
Anthropic itself also recommends prioritizing Sonnet 5.5 or Opus 5.5 for complex agent coding.
Haiku 5.5's value lies in the execution layer: tasks that have already been broken down, with clear acceptance criteria, and that can be run in parallel.
For example, Cognition's Devin uses Opus 5.5 as the main model and Haiku 5.5 as the sub-agent; the FrontierCode combination scored 66.2%, higher than either model running alone.
If your existing business uses Haiku 4.5, this time you can't just swap in a new model name and go live.
The budget_tokens manual thinking config now throws an error directly; it must be changed to adaptive thinking plus the effort parameter.
temperature, top_p, and top_k are all locked to default values; applications that rely on sampling parameters for creativity control or diverse generation need to change their logic.
Assistant message prefilling has been removed; the old approach of using prefilling to force a JSON opening must switch to tool calls or the structured output interface.
The computer use tool version was also updated, from computer_20250124 to computer_toolset_20260801; the interface format and return structure may both be different.
Also, adaptive thinking is enabled by default, so the first content block in a response may be a thinking block rather than the body; parsing logic needs to filter by the type field.
Five changes, and each one can make a request fail outright or return an unexpected structure.
Anthropic has provided a migration guide; it's recommended to go through it before switching.
Two perks along the way
Sonnet 5.5's cache reads dropped from $0.20/M to $0.10/M, and Anthropic says most agent tasks will see costs fall by about 20% as a result.
Starting this week, Max and Team subscribers can claim API credits: Max 5x gets $100 per month, Max 20x gets $200 per month, and Team gets up to $500/month shared within the team.
These credits can be used to experiment with the API and build tools, apps, and agents, and they work with all models.
What? You're saying Company A is giving me back the money I paid for my Coding Plan, and I get to use it once more on the API?
So what is OpenAI doing? OpenAI sent another reset card.
Reference links:
[1]The link https://www.anthropic.com/claude-haiku-5-5
[2]https://x.com/AI_Screening/status/2107906044915749274?s=20