AI Model Report

Reviews · OCTOBER 8, 2026

Haiku 5.5 Lands at $0.10/M Input, 1M Context, and a First-Ever Effort Dial

Anthropic cut Haiku's short-prompt token prices 90% on October 7, extended the context window to 1M tokens, and shipped an adjustable effort setting — the first on a Haiku-class model.

By Karl Strauchman · Senior model reviewer · October 8, 2026

Anthropic shipped Claude Haiku 5.5 on October 7, cutting input tokens to $0.10 per million and output tokens to $0.50 per million for prompts under 100,000 tokens, a 90% reduction against Haiku 4.5's $1.00/$5.00 and the exact pricing GPT-6 Luna already occupies. The release is also the first Haiku-class model with an adjustable effort parameter, and it ships with a 1M-token context window, a 128K max output, and 300K-token batch support in beta. Generally available day one on the Claude API, AWS, Google Cloud, and Microsoft Azure.

The pricing structure is the story under the story. Below 100,000 tokens, Haiku 5.5 is roughly a tenth of its predecessor. Above that threshold, prices step up 5× to $0.50 input and $2.50 output. Anthropic says about 90% of Haiku 4.5 requests fell under the 100K line, which is why the headline economics work; it's also why operators planning long-context workloads need to model the cliff carefully. Anthropic's own estimate is a ~75% average workload cost reduction below Haiku 4.5, which bakes in a new tokenizer that produces about 30% more tokens per unit of English text.

The benchmarks explain the confidence. On OSWorld 2.1's offline subset, Haiku 5.5 posts 72.4% against GPT-6 Luna's 48.9% and Haiku 4.5's 15.7%. On Terminal-Bench 4.0 it hits 39.2% at maximum effort and roughly 20% at medium, versus 16.4% for Luna and a flat 0.0% for Haiku 4.5. The effort dial is doing real work here, and it's the mechanism small teams will actually tune.

Partner numbers point the same direction. Asana reports up to 2.5× faster inference per agent turn and over 30% lower task-completion latency against its previous model. Vinh said Asana "saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn" compared with its current model. HubSpot's simulated CRM task suite averaged 92.8% across three runs under Haiku 5.5, according to Ze'ev Klapow, distinguished software engineer at HubSpot, who said it was the best score HubSpot had seen on that suite. AlphaSense, running Ask in Document across roughly 8 million calls a week, scored Haiku 5.5 at 0.84 against Haiku 4.5's 0.76 on 400 production-style queries. Daniel Campos, distinguished engineer at AlphaSense, called it a statistically significant improvement over Haiku 4.5 across the 400-query evaluation.

Anthropic simultaneously halved Sonnet 5.5's cache read price from $0.20 to $0.10 per million, which it estimates trims ~20% off typical agentic workloads. Max 5x and Max 20x subscribers now get $100 and $200 in monthly API credits; Team plans pool up to $500. Knowledge cutoff is June 2026.

The 90% price match to GPT-6 Luna isn't coincidence. It's the small-model floor becoming a commodity, with effort dials and context windows as the remaining axes of differentiation.

Sources