AI Model Report

Infrastructure · AUGUST 16, 2026

OpenAI's Ultrafast tier puts GPT-5.6 Sol at 750 tokens/sec on Cerebras wafers

A limited API preview launched August 13 runs OpenAI's flagship at 14× Standard throughput by keeping model weights entirely in the Wafer-Scale Engine's 44 GB of on-chip SRAM.

By Aiko Tanaka · Inference & serving · August 16, 2026

OpenAI opened a limited API preview of a new service tier called Ultrafast on August 13, serving GPT-5.6 Sol at up to 750 output tokens per second, or roughly 14× the throughput of its Standard tier. The mode runs entirely on Cerebras Wafer-Scale Engine hardware, and it's the first production surface of the ten-billion-dollar partnership the two companies signed earlier this year.

The mechanism is worth stating plainly, because it reframes what a frontier tier is competing on. Each Cerebras wafer carries 44 GB of on-chip SRAM, enough to hold the model weights themselves on the die. That eliminates the memory-bandwidth trip to HBM that governs how fast any GPU-based inference stack can emit tokens. OpenAI says the tier serves the same GPT-5.6 Sol weights as Standard, with no distillation and no capability tradeoff. The delta is pure plumbing.

The benchmark numbers Cerebras published alongside the launch are aggressive. On Humanity's Last Exam, a 2,500-question set spanning graduate-level chemistry, economics, and literature, GPT-5.6 Sol on Ultrafast finishes in just over 11 hours; Anthropic's Claude Fable 5 needs more than three days of continuous compute at comparable accuracy. That's a roughly 7× wall-clock advantage. On GDP-Val, OpenAI's benchmark of economically valuable knowledge-work tasks like legal briefs, financial models, and engineering reports, Cerebras reports a 5.6× end-to-end speedup "with no loss in quality." Using Artificial Analysis output-speed figures, Cerebras positions the tier as 5× faster than Claude Opus 4.8 Fast and 11× faster than Claude Fable 5.

Three details in the small print matter more than the headline multiple.

Pricing is unpublished, per TechCrunch and TheNextWeb, and there's no general-availability date. Access is gated to a small group of trusted partners whose participation OpenAI disclosed to the U.S. government ahead of release. And Ultrafast sits a third rung above the existing Fast Mode, which The Decoder notes offers about 2.5× speed at roughly double the Standard price. Whatever multiple Ultrafast lands on will price the entire premium-inference ladder.

The commercial choreography rewards a second look. Cerebras (NASDAQ: CBRS) went public earlier in 2026, then signed the ten-billion-dollar OpenAI deal, and now anchors the first tier where a frontier lab has publicly conceded that GPU inference isn't the ceiling of the product. Ultrafast is a benchmark launch as much as a service launch, and the party being benchmarked against by name is Anthropic. The last time a hardware-software pairing was staged this cleanly against a specific rival was the CUDA-era co-marketing that hardened the Nvidia-frontier-lab alliance in the late 2010s. The wafer, this time, is doing the talking.

Sources