AI Model Report

Reviews · SEPTEMBER 26, 2026

Reviewed: Claude Opus 5.5 at $4/$20, With a 60% Cache-Read Cut

Anthropic's new flagship posts 66.4% on Terminal-Bench 4.0, drops cache reads from $0.50 to $0.20 per million tokens, and becomes the default across Claude Code, Claude.ai, and Cowork on launch day.

By Karl Strauchman · Senior model reviewer · September 26, 2026

Anthropic shipped Claude Opus 5.5 on September 22, 2026, at $4 per million input tokens and $20 per million output, a 20% cut on both sides of the meter versus Opus 5. The cache-read line moved further: from $0.50 to $0.20 per million tokens, a 60% reduction that lands squarely on the workloads that reuse long context. Cache writes drop from $6.25 to $5. Anthropic's own math puts the blended effect around 40% cheaper than Opus 5 on typical workloads at default settings, alongside the same afternoon's broader frontier-price reset across OpenAI and Anthropic.

The benchmark card is unusually clean. Opus 5.5 posts 66.4% on Terminal-Bench 4.0 at xhigh effort (±2.6 pts), against 57.9% for GPT-6 Astra, 55.8% for Fable 5.1, and 52.3% for its own predecessor. It leads CursorBench 4.0 by 11 points over GPT-5.6 Sol, and pulls a 300+ Elo gap over OpenAI's flagship on GDPval-AA v2.1's 44-occupation knowledge-work suite (1846 vs. 1542; Fable 5.1 sits at 1735). On Humanity's Last Exam with tools it hits 67.7%. FrontierCode v1.1 lands at 54.4%, OSWorld 2.0 at 81.8%, Chartography at 89.0%. Artificial Analysis puts it at 58 on the Intelligence Index at max effort, leading six of ten core evals. It doesn't sweep everything: GPT-6 Astra takes Terminal-Bench-Science 0.1 (64.6% vs. 58.7%) and edges AutomationBench (41.4% vs. 40.0%).

The efficiency numbers are what earn the pricing. Output generation is more than 30% faster than Opus 5. The context window is 1M tokens, with up to 128,000 output tokens synchronously and 300,000 via the Message Batches beta. Anthropic's own HAProxy C-to-Rust port finished in 9.5 hours against 12 for Fable 5.1, at 51% lower cost. One early tester audited a 200,000-line codebase in under three hours; Opus 5 took over 20 and burned 2.5x the tokens. Another migrated 680,000 lines in under a day.

Customer statements track the same pattern. Kiro's VP of Agentic AI Deepak Singh reports 40% fewer calls and half the tokens versus Opus 5. Box VP of AI Products Yashodha Bhavnani says Opus 5.5 uses one-third of the tokens Opus 5 required and produces 40% less verbose output. Optiver's Noyan Tokgozoglu cites a 40–50% reduction in trading-support workload costs. Vellum characterizes the aggregate as a 25–40% decrease in raw token usage per task, with 5-hour rate limits raised 20%; Vellum adds that the reduced token pricing lets rate limits stretch 25% further.

A fast mode at $8/$40 offers up to 2.5x speed for latency-sensitive work. Safety telemetry, drawn from a ~2,000-scenario behavioral audit, shows Opus 5.5 attempting to circumvent boundaries about 85% less often than Opus 5 or Mythos 5.1.

The read: Anthropic has stopped selling a leaderboard and started selling a cost curve. That's the posture of a vendor that thinks agentic workloads, not single queries, are where the next two years of revenue live.

Sources

  • https://www.anthropic.com/claude-opus-5-5
  • https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/
  • https://venturebeat.com/technology/anthropic-releases-claude-opus-5-5-beating-fable-5-1-on-key-agentic-benchmarks-at-60-cheaper-api-price
  • https://www.vellum.ai/blog/claude-opus-5-5-benchmarks-explained
  • https://www.marktechpost.com/2026/09/22/anthropic-claude-opus-5-5-release/