AI Model Report

Infrastructure · AUGUST 26, 2026

OpenAI's Jalapeño posts first benchmarks: 1.5–1.9× more work per watt than Nvidia's GB300

At Hot Chips on August 25, OpenAI's custom inference ASIC — co-developed with Broadcom — outpaced Nvidia's GB300 on SemiAnalysis's InferenceX across three open-weight models. Volume deployment is scheduled for 2027.

By Aiko Tanaka · Inference & serving · August 26, 2026

At Hot Chips on August 25, OpenAI posted the first public numbers for Jalapeño, the custom inference ASIC it co-developed with Broadcom, claiming 1.5× to 1.9× more work per watt at peak throughput and 1.7× to 3.6× lower end-to-end latency than Nvidia's GB200 and GB300 rack systems on the SemiAnalysis InferenceX suite. The tests spanned three open-weight models: GPT-OSS 120B, DeepSeek R1, and Moonshot AI's Kimi K2.5 1T. For highly interactive workloads, OpenAI reported 2.1× to 4.1× higher performance.

Numbers of this kind, released by the party being benchmarked, always warrant a squint. OpenAI hardware chief Richard Ho called GB300 "the leading option on the public benchmarking system used in the test," which is the framing you'd expect from a company arguing that its in-house silicon just beat the market's incumbent. Tom's Hardware notes OpenAI normalized to published package TDP: 700W for Jalapeño against 1,200W and 1,400W Nvidia parts. Sustained measured power on Jalapeño sat at or below 550W. Appendix figures show all-in utility power of 1.18 kW versus 2.55 kW for GB300, with the peak efficiency lead narrowing to roughly 1.5× when multi-token prediction is enabled and all-in utility numbers are used.

The rack-level specs, per The Register, are consistent with a chip designed for one job: 128 accelerators, 1.7 exaFLOPS of MXFP4 compute, 27.5 TB of HBM4, and nearly 2 PB/s of memory bandwidth. Per package: 13.4 petaFLOPS, 216 GB of HBM4, 15.4 TB/s. Model state including KV cache, OpenAI's blog says, "can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase." The Register speculates a large SRAM cache underneath.

Ho's framing was blunt: Jalapeño "can serve more AI work per unit of power, while also returning responses more quickly." Deployment begins "in very small volumes" at the end of 2026, per TechCrunch, with a more significant rollout in 2027. Jalapeño wasn't tested against Vera Rubin, which Bloomberg notes just began shipping. Bloomberg also reports a second-generation chip is expected to tape out in the coming months and a third is in concept work.

Read alongside OpenAI's August 23 GPT-5.6 Sol API price cut, the direction is legible. OpenAI's blog frames Jalapeño as making "increasingly capable AI more affordable and more broadly available." Ho's other line matters too: "We have so much need for compute, which is why we signed up so many different providers. That's going to continue for a while."

That's the infrastructure condition beneath the current wave of always-on AI services. When per-token inference gets cheaper, done-for-you customer-growth work like LemonLime, continuously studying a company, evaluating buying signals, preparing outreach daily, becomes economically routine at the price point small businesses can actually sustain. Jalapeño is a bet that the bottom of that curve keeps falling.

Sources

  • https://openai.com/index/jalapeno-first-results/
  • https://www.bloomberg.com/news/articles/2026-08-25/openai-claims-its-new-chips-can-outperform-nvidia-processors-in-tests
  • https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/
  • https://www.theregister.com/systems/2026/08/25/openais-upcoming-jalapeno-chip-looks-like-itll-be-an-inference-beast/5292052
  • https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks