Model Releases · OCTOBER 9, 2026
Anthropic's Claude Haiku 5.5 lands at $0.10/M under 100K, 75% cheaper on average than Haiku 4.5
The October 7 release pairs a 90% nominal input-token cut with the first Haiku-class effort control, a Sonnet 5.5 cache-read halving, and new monthly API credits for Max and Team subscribers.
Anthropic released claude-haiku-5-5 on October 7, pricing input tokens at $0.10 per million and output at $0.50 per million for any prompt under 100,000 tokens. That's a 90% nominal cut against Haiku 4.5's $1.00 input / $5.00 output. It also happens to match the headline rate OpenAI set for GPT-6 Luna, which is almost certainly the point.
The pricing cliff matters more than the headline. Above 100K tokens, Haiku 5.5 jumps to $0.50 input and $2.50 output, five times its own lower tier, and what Anthropic frames as a 50% cut from Haiku 4.5. Anthropic says roughly 90% of Haiku 4.5 requests landed below 100K, so for most operators the effective savings are real. Factoring in an updated tokenizer that consumes somewhat more tokens per unit of work, the company puts realistic average workload savings at 75%. Luna, by comparison, doesn't surcharge until 272,000 input tokens, which is where the two pricing cards stop matching.
What's quietly more interesting is the capability jump at the small-model tier. Haiku 5.5 posts 72.4% on OSWorld 2.1's offline subset against Haiku 4.5's 15.7%, and 45.9% on Humanity's Last Exam without tools versus 10.2%. On GDPval-AA v2.1, it scores 1,620 to Haiku 4.5's 735. The Terminal-Bench 4.0 headline figure is 39.2%, but that's at maximum effort; at medium, the default, it's about 20%. VentureBeat flagged the gap, and it's a fair flag, the marketed numbers aren't what shows up at default settings.
Haiku 5.5 is also Anthropic's first Haiku-class model with an adjustable effort setting, and its first to ship with built-in safeguards for a narrow band of high-risk cybersecurity requests, per Reuters. Anthropic says most everyday tasks are unaffected.
Alongside the release, Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens, Anthropic estimates about 20% savings on typical agent workloads, stacking on the 30% speed improvement and up to 30% per-task cost cuts that landed with Sonnet 5.5 in late September. Max 5x subscribers now get $100 in monthly API credits, Max 20x gets $200, and Team subscriptions get up to $500 shared.
Asana's Aaron Vinh, in Anthropic's launch materials, reports over 30% latency reduction and up to 2.5x faster inference per agent turn, which, as vendor-supplied figures tend to be, is directional rather than independent.
The through-line connecting this release to the simultaneous Opus 5.5 and GPT-6 Sol/Luna drops two weeks earlier is price matching at every tier, not performance leadership. The frontier labs are converging on identical rate cards ahead of Anthropic's planned IPO. The small model got cheap because it had to.
Sources
- https://www.anthropic.com/claude-haiku-5-5
- https://aws.amazon.com/blogs/machine-learning/introducing-claude-haiku-5-5-on-aws/
- https://venturebeat.com/technology/anthropic-launches-claude-haiku-5-5-with-90-api-price-reduction-matching-gpt-6-luna
- https://www.marktechpost.com/2026/10/07/anthropic-releases-claude-haiku-5-5-a-small-model-with-1m-context-priced-at-0-10-per-million-input-tokens/
- https://money.usnews.com/investing/news/articles/2026-10-07/anthropic-launches-third-claude-5-5-model-expanding-ai-lineup-before-planned-ipo