Infrastructure · SEPTEMBER 12, 2026
Sakana Ships Fugu Max at $2/$6 per Million Tokens, Splitting Orchestration Into Cost and Capability Tiers
Fugu Max routes across a swappable pool of open-weight and specialized models — including NVIDIA Nemotron — behind one OpenAI-compatible API, with output priced 40–60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3.
Sakana AI split its Fugu line into two tiers on September 11, pricing the new Fugu Max at $2 per million input tokens and $6 per million output tokens, roughly 40–60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3 on output. Fugu Ultra v2, positioned as the frontier-adjacent option, lists at $5/$30 per million with a jump to $10/$45 above 272,000 tokens. Both share a 1,000,000-token context window and a 128,000-token maximum output. Ultra v2's training cutoff is August 28, 2026.
The Tokyo-based lab isn't shipping a single model. Fugu is an orchestrator that routes each request across a swappable pool of open-weight and specialized models, including NVIDIA's Nemotron family through a partnership announced in August, all sitting behind one OpenAI-compatible API. Sakana cites two ICLR 2026 papers, TRINITY: An Evolved LLM Coordinator and Learning to Orchestrate Agents in Natural Language with the Conductor, as the intellectual scaffolding, and points to Vercel AI Gateway as an access channel. A security-focused variant, Fugu Cyber, is reported at 86.9% on CyberGym and 72.1% on CTI-REALM.
The pricing is the story. This is what a serious cost-tier attack on Anthropic and OpenAI's inference economics looks like when it lands in the same week as a full DeepSeek architecture roll, and buyers of AI-assisted outreach, qualification, and chat features now have a fresh floor to negotiate against. Cached input on Fugu Max lists at $0.25 per million tokens on OpenRouter's provider table. Ultra v2's long-context cached input starts at $0.50.
The benchmark story is louder but softer. Sakana claims Fugu Max takes best-overall on six evaluations spanning Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and the internal SWEFish, and extends the cost-performance Pareto frontier on seven of ten. Ultra v2 claims top-two on seven of eight benchmarks and best or joint-best on five, including a 74.3 on DeepSWE and a 48.3 on Chartography against Opus 5's 27.3 and Fable 5's 29.5, roughly 77% higher, per Pondero's math. Sakana says Ultra v2 outperforms models costing three to five times more per token.
DataNorth notes that only two benchmark numbers appear as text in Sakana's announcement; the rest live inside chart images, and SWEFish is Sakana's own unreleased suite. OpenRouter's outside measurement clocks Fugu Max at 13 tokens per second with a 5.37-second first-response. Sakana's own FAQ concedes buyers can't see which model answered a given prompt, and that the pool can change without notice. Fable 5, Fable 5.1, and GPT-6-Astra are explicitly excluded from the Ultra v2 agent pool.
Routing opacity is a feature to Sakana and a procurement problem for anyone whose compliance stack expects to know which weights produced which token. The $6 output price will move the market first. The orchestration claims will wait for reproduction.
Sources
- https://sakana.ai/fugu-max-release/
- https://datanorth.ai/news/sakana-ai-launches-fugu-max-and-fugu-ultra-v2
- https://theroboticsmedia.com/article/sakana-ai-fugu-ultra-v2-0-1m-context-multi-agent-orchestration-september-11-2026
- https://pondero.ai/news/2026-09-12-sakana-fugu-max-ultra-v2/
- https://alphasignal.ai/news/sakana-ai-splits-fugu-into-max-and-ultra-v2-to-cut-costs-60