Reviews · SEPTEMBER 29, 2026
Claude Sonnet 5.5 ships at Sonnet 5 pricing, jumps Terminal-Bench 4.0 to 70.6%
Anthropic held the $2/$10 token price and pushed per-task cost down as much as 30% by cutting tokens, tool calls, and steps rather than the sticker rate.
Anthropic released Claude Sonnet 5.5 on September 28, 2026 at the same $2 per million input tokens and $10 per million output tokens that Sonnet 5 has carried since June, with cache reads at $0.20 per million and cache writes at $2.50 per million. The sticker didn't move. The economics did.
Anthropic's own testing puts output generation more than 30% faster and per-task cost as much as 30% lower, because the model finishes agentic work with fewer tokens, fewer tool calls, and fewer steps. That's the whole trick. Held pricing, compressed work.
The benchmark that tells the story is Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% against Sonnet 5's 10.3% and Opus 5.5's 66.4% at Xhigh effort. A mid-tier model at half the price of Opus ($4 input, $20 output, $5 cache writes) is now finishing agentic terminal tasks its predecessor couldn't complete at all. On CursorBench 4.0, Sonnet 5.5 lands at 55.5% against Opus 5.5's 57.8%. On FrontierCode 1.1's main set, Sonnet 5.5 reaches 46.2% at Max effort and 52.1% at Xhigh, up from Sonnet 5's 42.4% and inside striking distance of Opus 5.5 at 54.4% and GPT-6 Sol at 49.3%. GDPval-AA v2.1 moves from 1449 to 1844, essentially matching Opus 5.5's 1846. OSWorld 2.1 partial credit lands at 80.1%; Humanity's Last Exam with tools at 64.5%; Chartography without tools at 61.6%.
The customer testimonials are unusually consistent in what they measure. Yashodha Bhavnani, VP of AI Products at Box, reports the model ran "2.4x faster and used 12% fewer total tokens." Curtis Allen, principal engineer at Slack, saw roughly 14% fewer output tokens on offline Slackbot evaluations with no prompt changes. Abhinay Kathuria, director of AI at Zendesk, cites 20% faster ticket processing against Claude models currently in production. Atlassian says its Rovo Agents run up to 30% faster on Sonnet 5.5 versus Sonnet 5.
The compression numbers get sharper at the workflow layer. Balyasny Asset Management, running a set of 2,441 finance tasks, saw token consumption fall from 497,000 per answer on Sonnet 5 to roughly 121,000 on Sonnet 5.5. Fabian Hedin, CTO at Lovable, reports about one-third fewer tool calls and roughly half as many shell executions on coding jobs. Gabriel Grinberg at Base44 measured 3.6 average iterations per build across 118 real applications, against 7.7 with Opus 5.
There's a pricing history worth remembering here. Anthropic introduced Sonnet 5 at $2/$10 as introductory pricing in June, with a previously planned move to $3/$15 that it quietly abandoned in August. Sonnet 5.5 arriving at the frozen price completes what was, in retrospect, a soft repricing of the tier. The AutomationBench results placing Opus 5.5 at 40% on real business workflows already reframed how much the top tier was actually worth; Sonnet 5.5 answers the question from below.
The interesting number on this model is what it costs to finish the job.
Sources
- https://www.anthropic.com/claude-sonnet-5-5
- https://techcrunch.com/2026/09/28/anthropic-releases-sonnet-5-5-which-it-calls-a-significantly-cheaper-faster-work-partner/
- https://github.blog/changelog/2026-09-28-claude-sonnet-5-5-in-github-copilot/
- https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-5-with-30-cost-reduction-per-task-due-to-faster-speeds-and-fewer-tool-calls
- https://www.unite.ai/anthropic-releases-claude-sonnet-5-5-at-unchanged-sonnet-5-pricing/