AI & Machine Learning
·By Seedwire Editorial·

Claude Haiku 5.5 cuts prices, with a prompt-length catch

Claude Haiku 5.5 cuts prices, with a prompt-length catch

Illustration, not documentary evidence of the event.

Anthropic has released Claude Haiku 5.5, a small model aimed at high-volume tasks such as summarization, classification and customer support. According to the-decoder.com, reporting Anthropic’s announcement, the model adds adjustable reasoning levels and is available through Amazon Web Services, Google Cloud and Microsoft Azure.

The reported pricing makes prompt length a consequential purchasing detail. For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, compared with $1.00 and $5.00 for Haiku 4.5. Above that prompt threshold, Haiku 5.5’s rates rise to $0.50 for input and $2.50 for output. Anthropic also says its updated tokenizer consumes slightly more tokens per task, so the token-price reduction should not be treated as an equivalent reduction in a workload’s bill.

What the threshold costs in a worked example

Using Anthropic’s published input and output rates, assume 1,000 uncached requests, each with 100,000 input tokens and 1,000 output tokens. That is 100 million input tokens and one million output tokens: (100 × $0.10) + (1 × $0.50) = $10.50.

Increase each prompt by just one token, to 100,001, while holding output constant. The higher rate now applies: (100.001 × $0.50) + (1 × $2.50) = $52.5005, approximately $52.50. In this hypothetical batch, 1,000 additional input tokens increase the bill by about $42 because the requests cross a pricing boundary, not because those tokens alone are expensive.

This calculation isolates the threshold. It is not a benchmark, observed customer bill or savings forecast. It excludes caching, discounts, tools, retries and changes in generated output. Token counts are assumed, not measured; count the complete request under the new tokenizer before applying the example. For an application close to the boundary, a small context change can matter more to cost than its size suggests. Removing essential context merely to stay below it could reduce answer quality.

Anthropic’s reported benchmarks support evaluating Haiku for more work, but do not establish that it can replace a larger model across an application. On Terminal-Bench 4.0, its agentic coding score is 39.2 percent, against Sonnet 5.5’s 70.6 percent. Anthropic itself recommends narrowly scoped work for Haiku and larger models for complex agentic coding. A practical evaluation would therefore separate summarization or classification from tasks requiring extended coding work, with acceptance criteria for each rather than a single application-wide score.

Existing Sonnet users have a separate cost decision. Anthropic is cutting Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens and estimates roughly 20 percent savings for most agentic tasks. Before changing models, inspect how much of the current bill comes from cache reads and apply the new rate to that usage. That gives teams a revised Sonnet baseline against which to judge whether Haiku’s additional savings justify any measured quality tradeoff.

Claude Haiku 5.5
Anthropic
API pricing
prompt caching
model evaluation
Seedwire Newsletter

Follow Seedwire by email

Request Seedwire news emails. There is no guaranteed delivery schedule. You can withdraw your request through the privacy contact.

By selecting Subscribe, you request Seedwire news emails. Privacy and withdrawal.