Model Releases · OCTOBER 9, 2026
Haiku 5.5 at $0.10/M: the small-model tier just crossed the price threshold for live customer workflows
Anthropic's October 7 release cuts Haiku input pricing 90% for prompts under 100K tokens, lifts Terminal-Bench 4.0 from 0% to 39.2%, and posts 72.4% on OSWorld 2.1 — reshaping what a small team can afford to run against live customer traffic.
Anthropic released Claude Haiku 5.5 on October 7 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, a 90% cut on the input side against Haiku 4.5's $1 rate (and a drop from $5 to $0.50 on output). Above the 100K threshold, the pricing shifts to $0.50 input and $2.50 output, roughly five times more. Anthropic says 90% of past Haiku requests fell under the threshold, and estimates a 75% blended average cost reduction once a new tokenizer is factored in.
That's the price story. The capability story is why it matters.
Haiku 4.5 scored 0.0% on Terminal-Bench 4.0. Haiku 5.5 posts 39.2%. On OSWorld 2.1's offline subset the jump is 15.7% to 72.4%, and GDPval-AA v2.1 moves from 735 to 1,620. Sonnet 5.5 still leads on Terminal-Bench at 70.6%, but Haiku 5.5 is now in the conversation rather than the demo reel. On Vals.ai it ranks #16 of 45 at 54.31%, and #3 of 110 on Vibe Code Bench at 90.44%. Artificial Analysis clocks it at 43 on their Intelligence Index.
The comparison that'll animate buyer conversations is OpenAI's GPT-6 Luna, priced identically at $0.10/$0.50. Haiku 5.5 beats it 72.4% to 48.9% on OSWorld 2.1 and 39.2% to 16.4% on Terminal-Bench 4.0, with a 40% hallucination rate against Luna's 77% on Artificial Analysis's test. The caveat: Haiku burns about 162,000 output tokens per task at highest effort versus Luna's ~50,000, roughly 3x. At parity pricing, verbosity becomes the real cost center.
For a small team, the arithmetic has shifted. A lead-classification pipeline chewing a million input tokens a day runs about $0.10 on Haiku 5.5's short-prompt tier versus $2 on Sonnet 5.5 at $2/$10, a twenty-fold gap before output is counted. That makes live chat triage, CRM audit passes, outreach subagents, and signal-parsing loops newly affordable at production volume rather than toy scale.
External validation is beginning to arrive. Aaron Vinh, staff software engineer at Asana, reports over 30% latency reduction and up to 2.5x faster inference per agent turn. AlphaSense tested the model against roughly 8 million feature calls a week. Rogo is in the mix. Anthropic also halved Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens, which it estimates will cut most agentic Sonnet workloads by about 20%. The model carries a 1M context window and ships across AWS, Google Cloud, and Microsoft Azure.
Weak spots remain legible: 5.73% on SRE Bench, 1.25% on Harvey's Legal Agent Benchmark. Haiku 5.5 isn't a specialist. It's a volume layer, and the volume layer just got cheap enough to point at live customers.
Sources
- https://www.anthropic.com/claude-haiku-5-5
- https://the-decoder.com/claude-haiku-5-5-arrives-with-massive-price-cuts-proving-the-ai-pricing-arms-race-is-far-from-over/
- https://siliconangle.com/2026/10/07/anthropic-releases-claude-haiku-5-5-small-model-and-halves-sonnet-5-5-cache-read-prices/
- https://www.vals.ai/models/anthropic_claude-haiku-5-5
- https://www.marktechpost.com/2026/10/07/anthropic-releases-claude-haiku-5-5-a-small-model-with-1m-context-priced-at-0-10-per-million-input-tokens/