AI Model Report

Infrastructure · 6 pieces on file

Infrastructure

Serving stacks, kernels, GPU economics, and the gap between published throughput and what reproduces on real hardware.


Feature · AUGUST 16, 2026

OpenAI's Ultrafast tier puts GPT-5.6 Sol at 750 tokens/sec on Cerebras wafers

A limited API preview launched August 13 runs OpenAI's flagship at 14× Standard throughput by keeping model weights entirely in the Wafer-Scale Engine's 44 GB of on-chip SRAM.

By Aiko Tanaka · Inference & serving

Read the full piece →


More in Infrastructure