雷峰网 AI

Behind the cost cuts of Claude Haiku 5.5: Computer-use capabilities soared, but why is complex coding still falling short?

The model now dynamically allocates reasoning by task; beyond 100,000 Tokens, both input and output unit prices rise 5x. Author | Zheng Jiamei Editor | Cen Feng Just now, Anthropic released Claude Haiku 5.5. The upgrade…

The model now dynamically allocates reasoning by task; beyond 100,000 Tokens, both input and output unit prices rise 5x.

    Author | Zheng Jiamei

Editor | Cen Feng

                                                                                                       

Just now, Anthropic released Claude Haiku 5.5. The upgrade is substantial, especially in computer use, where the OSWorld 2.1 benchmark score rose from the previous generation's 15.7% to 72.4%. Knowledge work and complex reasoning capabilities have also improved markedly, and according to data published by Anthropic, the new model's average operating cost has dropped by about 75%. In addition, Haiku 5.5 adds an adaptive reasoning mechanism and supports a 1 million Token context window. In the past, Haiku was mainly used for high-frequency tasks like summarization, classification, and information extraction, with its advantages always being speed and low cost. After this update, the work it can handle has become considerably more complex, with some benchmark scores even approaching Sonnet 5.5. However, on complex coding tasks, the gap between the two remains clear. Looking at the updated reasoning mechanism and pricing rules, there are actually quite a few details to unpack regarding Haiku 5.5's performance across different tasks and its real usage costs.

01


Complex coding still lags

Haiku 5.5's performance gains are mainly reflected in computer use, knowledge work, and multidisciplinary reasoning, with the change in computer use being especially pronounced. Leiphone (public account: Leiphone) On the OSWorld 2.1 offline subset, Haiku 5.5 scored 72.4%, while the previous generation only managed 15.7%, and Sonnet 5.5 reached 83.9%. OSWorld tests a model's ability to complete long-chain tasks through a computer interface, involving interface understanding, action execution, and continuing tasks based on environmental feedback. Compared with ordinary text Q&A, computer use involves more state dependency. After the model completes a click or input, the page may change, and subsequent actions must be decided based on the new interface state. If an operation fails, the model must also judge whether the current environment still meets the task requirements. Haiku 5.5's growth on this benchmark shows that its ability to handle sequential operation tasks has improved markedly over the previous generation. But 72.4% corresponds to the completion rate on a specific test set and cannot directly represent the actual success rate across all browser or desktop tasks. Dynamic pages, permission restrictions, and abnormal states in real business scenarios still need to be tested separately in the corresponding environments. As for knowledge work, Haiku 5.5 scored 1620 points on GDPval-AA v2.1, Haiku 4.5 scored 735, and Sonnet 5.5 scored 1840. This benchmark covers real work tasks across 44 categories of occupations, involving material analysis, information integration, and work-product generation. In another benchmark, AA-Briefcase v1.1, Haiku 5.5 improved from the previous generation's 614 points to 1578 points. On the multidisciplinary reasoning test Humanity's Last Exam, the no-tools score reached 45.9%, and 57.4% with tools, also clearly higher than the previous generation. The tasks involved in these benchmarks are more complex than ordinary information extraction. The model needs to understand the relationships among multiple information sources, handle constraints in the materials, and form results based on existing content. The introduction of tools also adds steps for information retrieval and result integration. However, the coding benchmarks present a different picture. On the FrontierCode 1.1 main test, Haiku 5.5 scored 46.4%, a small gap from Sonnet 5.5's 52.1%. But on Terminal-Bench 4.0, Haiku 5.5 scored 39.2%, while Sonnet 5.5 reached 70.6%. Terminal-Bench involves complex multi-step tasks in a command-line environment, where the model typically needs to read files, execute commands, handle errors, and adjust operations based on test feedback. It places high demands on sustained planning and execution-state maintenance. Anthropic also explicitly stated in its release materials that Sonnet 5.5 and Opus 5.5 remain better suited for complex agentic coding tasks, while Haiku 5.5 is mainly aimed at clearly scoped subtasks, summarization, and context compression work. Early enterprise testing provided more specific data. Asana reported that in AI Teammates-related tests, Haiku 5.5 reduced task completion latency by more than 30% compared with its current model, with single-turn inference speed increasing by up to 2.5x. HubSpot measured an average score of 92.8% across three runs in a simulated CRM environment, while AlphaSense measured a score of 0.84 for Haiku 5.5 on 400 document Q&A tests, versus 0.76 for the previous generation. These results come from internal tests at different companies, each using different tasks and evaluation standards; they are useful for understanding performance in specific business scenarios but cannot serve directly as a unified model capability ranking.

02


Adaptive reasoning comes to Haiku

Beyond the performance changes, Haiku 5.5 is also the first Haiku model to support adjustable reasoning intensity. Leiphone Previously, when Haiku 4.5 used the extended thinking feature, developers had to set a fixed internal reasoning budget in advance. Whether the task was simple information lookup or complex analysis involving multiple conditions, the reasoning space available to the model was limited by the pre-set value. Haiku 5.5 switches to Adaptive Thinking. The model can decide whether to use internal reasoning and how much computation to invest based on the task content, while developers set the overall reasoning intensity via Effort. For example, extracting a date from a document usually does not require lengthy analysis, whereas comparing clause differences across multiple contract versions requires processing many more conditions and interrelated pieces of information at once. Adaptive reasoning allows the model to adopt different levels of computational investment for these two types of tasks. This mechanism also involves the allocation of inference-time compute. For complex tasks, increased internal reasoning lets the model handle more intermediate steps, but a longer reasoning process also increases Token consumption and response time. For simple tasks, reducing internal reasoning lowers unnecessary computational overhead. Anthropic showed on its release page the cost-performance relationship for Haiku 5.5 on OSWorld, GDPval-AA, and Humanity's Last Exam at different reasoning intensities. However, different tasks do not respond consistently to extra reasoning compute, so raising the reasoning level cannot be directly equated with a fixed magnitude of performance gain. There are also several changes at the API level to deal with. First, Haiku 5.5 no longer follows the previous generation's approach of manually specifying a fixed number of thinking Tokens, so existing reasoning configurations need adjustment. Second, internal reasoning consumes the model's output Token budget. If an existing program sets output limits based only on the length of the final answer, after the upgrade internal reasoning may take up too much space, preventing the body of the response from being fully generated. In addition, with adaptive reasoning enabled, model responses may contain separate thinking blocks. Applications need to correctly distinguish internal reasoning content from final output and can no longer assume the returned content necessarily begins with the answer. For agents that use tools continuously, there is also a consistency requirement between reasoning data and conversation history. If a developer modifies existing history messages in subsequent requests, it may affect the validity of the original reasoning state. Haiku 5.5 also adds restrictions on some sampling parameters, so applications that relied on these parameters to adjust output behavior need to check API compatibility. These changes mainly affect projects already using the Haiku 4.5 extended thinking feature. When migrating, besides swapping the model version, developers also need to re-examine reasoning budgets, response parsing, and multi-turn message handling.

03


No uniform low price for the million-token context

Haiku 5.5 supports a 1 million Token context and 128,000 Token output, capable of handling longer documents, large codebases, and continuously running tool logs. However, Anthropic did not set a uniform low price for the entire context window. For prompts of no more than 100,000 Tokens, input costs $0.10 per million Tokens and output costs $0.50 per million Tokens. Once a prompt exceeds 100,000 Tokens, the input price rises to $0.50 and the output price to $2.50. Caching fees are also tiered. The per-million-Token cache read price is $0.01 for shorter prompts and $0.05 for longer prompts. Compared with repeatedly passing in full content, caching is suitable for reusing system instructions, tool definitions, and shared reference material. Anthropic says about 90% of Haiku 4.5's requests are within 100,000 Tokens. These requests see a large price reduction under the new pricing, but actual costs are also affected by changes in Token usage. Haiku 5.5 has an updated tokenizer, so the Token count for the same content may differ from the previous generation. Therefore, even if the input text is unchanged, the actual billed quantity after upgrading may change and needs to be re-measured. For long-running agents, context management further affects costs. For example, a coding task may continuously produce file retrieval results, code analysis records, and test logs. If every call carries the full history, the input length keeps growing as the task progresses. Prompt Caching can reduce the read cost of repeated content, while Compaction can condense completed operations into a shorter task state, reducing the history information that subsequent calls need to process. Anthropic lists context compression as one of Haiku 5.5's main use cases in its release materials. Regarding specific deployments, Rogo reported a case of using Haiku subagents to extract revenue data from corporate 10-K filings, with a larger model handling the subsequent presentation creation. Cognition, meanwhile, uses Haiku 5.5 as an auxiliary model for Devin Fusion; in a configuration led by Opus 5.5, the FrontierCode score reached 66.2%, with reported reductions in cost and latency. These cases adopt an approach where different models handle different tasks, but the release materials do not provide complete call counts, Token distributions, or cost breakdowns, so it is impossible to calculate from them the savings a typical business would achieve with the same architecture. This time Anthropic also cut Sonnet 5.5's cache read price by 50%, from $0.20 per million Tokens to $0.10. According to Anthropic's figures, this reduces Sonnet 5.5's operating costs on most agentic tasks by about 20%. This adjustment only involves cache read fees; the standard prices for input and output Tokens have not been lowered by the same proportion.

04


Migration still requires checking real-world task performance

Claude Haiku 5.5 is now available through platforms such as the Claude API, AWS, Google Cloud, and Microsoft Azure, and Anthropic has also begun updating the Python and TypeScript SDKs, adding beta support for computer use and browser use. Based on the data released so far, Haiku 5.5 scores notably higher than the previous generation in computer use, knowledge work, and multi-disciplinary reasoning benchmarks, but it still lags Sonnet 5.5 by a considerable margin in complex terminal programming. In terms of pricing, shorter prompts enjoy a more pronounced cost advantage, while a different pricing tier applies once usage exceeds 100,000 tokens. For existing Haiku 4.5 applications, migration involves concrete changes such as inference configuration, token counting, and response parsing. If Sonnet was previously used for some high-frequency tasks, you can also compare Haiku 5.5's completion rate, response time, and invocation costs on the same test set. Anthropic has provided information on model capabilities, API changes, and pricing, but the actual cost improvements in real business still need to be validated against the corresponding workloads. Especially in multi-turn tool-calling scenarios, the model's per-request price, the number of retries after task failures, and context growth need to be accounted for within the same complete task.

Hop on board and let us show you the highlights of top global AI conferences

Exclusive access to:

Expert presentation slides

Full conference reports

Popular paper explainers

Interviews with rising academic stars

Scan the QR code above

or click "Read Original" to follow the channel.

This is an original Leifeng Network article; unauthorized reproduction is prohibited. For details, seeRepublishing Notice。

Original source

雷峰网 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original