AI & Machine Learning
·By Seedwire Editorial·

Opus 5.5 cuts prices 20%, but max-effort runs burn more tokens

Opus 5.5 cuts prices 20%, but max-effort runs burn more tokens

Illustration, not documentary evidence of the event.

Anthropic has released Claude Opus 5.5, the first model in a new family, and is selling it on two claims: it performs at the level of Claude Fable 5.1 on most tasks, and it costs roughly 40 percent less to operate than Opus 5, according to the-decoder.com, which based its report on Anthropic's announcement and on independent testing by Artificial Analysis. Sonnet 5.5 and Haiku 5.5 are due in the coming weeks.

The list price drops from $5 to $4 per million input tokens and from $25 to $20 per million output tokens, a 20 percent cut. Cache reads fall from $0.50 to $0.20, a 60 percent reduction, and cache writes go from $6.25 to $5. Anthropic says the model also uses fewer tokens and produces output more than 30 percent faster, which is how it gets from a 20 percent price cut to a 40 percent operating cost claim. Subscribers get a 20 percent bump in five-hour usage limits and can bank a limit reset for later. The model is live on Amazon Web Services, Google Cloud, Microsoft Azure, and the Claude Platform under the ID claude-opus-5-5.

On Anthropic's own table, Opus 5.5 leads Fable 5.1 and OpenAI's GPT-6 Astra on most rows, including 66.4 percent on Terminal-Bench 4.0 against 55.8 for Fable 5.1 and 57.9 for Astra. Astra keeps the edge on AutomationBench (41.4 versus 40.0) and Terminal-Bench-Science (64.6 versus 58.7). Anthropic frames the coding results as a cost story: it says Opus 5.5 beats Astra on FrontierCode at about a fifth of the cost per task and matches it on Terminal-Bench at 40 percent of the cost.

The 40 percent number depends on how hard you push the model

Artificial Analysis, the independent evaluation platform, puts Opus 5.5 at 58 on its Intelligence Index, the highest score it has recorded, with Fable 5.1 and GPT-6 Astra tied at 53. It also reports a caveat that complicates Anthropic's efficiency pitch. At maximum effort, Opus 5.5 consumed about 119,000 output tokens per task in its testing, compared with 73,000 for Opus 5, 78,000 for Fable 5.1, and 27,000 for Astra. The lower per-token prices bring cost per task back in line with Opus 5, but not below it. Four of the model's five effort levels do sit on the cost-performance Pareto frontier, so the cheaper runs are cheap for what they deliver.

This is the practical decision for anyone moving workloads over. The cost saving Anthropic advertises appears to come from running the model at lower effort settings than the ones that produce the headline benchmark scores. A team that reflexively selects max effort for everything may see the same bill as before, with a faster and more capable model but no savings. The Decoder also notes that Anthropic can no longer run this model with thinking mode disabled, which removes the cheapest configuration some integrations relied on. Effort selection becomes a real engineering choice rather than a default.

Safeguards that quietly swap the model underneath you

Opus 5.5 is the first Opus model to carry the same restrictions on cybersecurity, biology, and frontier LLM development that Fable 5.1 has. When a request trips those safeguards, Anthropic says it is routed to another model: most cybersecurity tasks go to the older Opus 4.8, while requests flagged for biology or frontier LLM work go to Opus 5. Ordinary bug fixing is meant to stay on Opus 5.5. Verified organizations can apply for biology access through the Life Sciences Verification Program, and the Cyber Verification Program is expected to cover Opus 5.5 within weeks.

Anthropic describes the routing as transparent, but the more interesting question is what that means in practice for someone building a security product or a pipeline that occasionally touches those domains. If a classifier misfires, the caller gets a two-generations-older model for that request, with different behavior and presumably different quality. Whether the API surfaces which model actually answered, and whether billing reflects it, is not addressed in the report. Teams in those fields should test the boundaries before assuming consistent output.

The other new constraints are aimed at competitors rather than customers. Anthropic says it is seeing industrial-scale distillation attacks, where thousands of fake accounts are used to extract a model's capabilities, and cites a September 2026 threat report on activity it has stopped. Its answer is Preserved Thinking, first shipped with Fable 5.1 and now applied to Opus 5.5, which stops API users from editing prior context to pull out the model's reasoning. It applies to API accounts created on or after August 31, 2026, so older accounts are grandfathered for now. Watermarking for EU AI Act compliance is included, as is a Zero Data Retention option.

Read together, the release suggests Anthropic feels pricing pressure from two directions, from OpenAI at the top and from cheaper Chinese models below, and is responding with a price cut, a promise of plainer writing, and a tighter perimeter around the model's internals. The promised fix for "Claudish" prose, meaning verbose, formulaic output, rests on early tester impressions rather than any benchmark in the report, so treat it as a claim until users report back. What to watch is whether Sonnet 5.5 and Haiku 5.5 arrive with the same effort-level tradeoff, because that is where most high-volume workloads actually run.

Claude Opus 5.5
Anthropic
Claude Fable 5.1
GPT-6 Astra
Artificial Analysis
LLM pricing
model distillation
EU AI Act
Seedwire Newsletter

Follow Seedwire by email

Request Seedwire news emails. There is no guaranteed delivery schedule. You can withdraw your request through the privacy contact.

By selecting Subscribe, you request Seedwire news emails. Privacy and withdrawal.