Model Releases · AUGUST 3, 2026
OpenAI cuts GPT-5.6 Luna 80% three weeks after launch, admits the pricing floor moved
Luna drops from $1/$6 to $0.20/$1.20 per million tokens, Terra falls 20%, and Sol gets a 2× Fast mode — OpenAI's fastest-ever post-launch reprice, pushed by Claude Opus 5, Gemini 3.6 Flash, and cheap Chinese open weights.
OpenAI cut the price of GPT-5.6 Luna by 80% on July 30, three weeks after shipping it, dropping the model from $1/$6 per million input/output tokens to $0.20/$1.20. Terra fell 20% in the same announcement, from $2.50/$15 to $2/$12. Sol's standard rate held at $5/$30, but a new 2× Fast mode at $10/$60 replaced the old Priority Processing tier and aligned with Codex's /fast.
Repricing a frontier model twenty-one days after launch isn't routine housekeeping. It's a vendor conceding, in public, that the floor moved underneath them between the spec sheet and the invoice.
The competitive geometry explains the speed. Anthropic shipped Claude Opus 5 at $5/$25, the same sticker as Opus 4.8, landing roughly 6% cheaper per token than Sol Standard at comparable performance. Google introduced Gemini 3.6 Flash at $1.50/$7.50 and Gemini 3.5 Flash-Lite at $0.30/$2.50 about ten days before OpenAI's move. Underneath all of that sit Kimi K3 from Moonshot, DeepSeek's flash tier, and Xiaomi's MiMo-V2.5 Flash, Chinese open weights that anchor the low end of every enterprise procurement conversation whether developers deploy them or not.
OpenAI's own framing is telling. The company attributes Luna's economics to a pipeline change that lifted prompt-cache reuse from 24% to 90%, letting the model handle 2.2× more context while emitting 8.5× fewer output tokens across production traffic. It positions Luna as 87% cheaper than GPT-5.4 mini, and cites Devin Fusion running 40% faster and 40% cheaper on the new default. That's a coherent efficiency story. It's also the story a vendor tells when the alternative is admitting that Gemini 3.6 Flash forced the reprice.
The Sol Fast tier is the more revealing move. At $10/$60 for up to 2.5× throughput at the same underlying intelligence, OpenAI is quietly rebuilding tiered inference as a product surface, replacing the API-side Priority Processing SKU with something branded and consumer-facing. Sticker prices go down at the bottom; latency becomes the thing you pay for at the top.
The lineup now spans a 50× band from Luna's $1.40 combined to Sol Fast's $70. A single vendor covering that range is what happens when the frontier stops being a scarcity story. Fable 5, Opus 5, Gemini 3.1 Pro, and the Chinese flash tier all price into different corners of the same grid, and buyers are learning to route by cost per task rather than loyalty to a house.
The three-week interval is the artifact worth keeping. OpenAI's previous cadence gave a model months to earn its keep. Now the market gives it weeks, and the pricing power that used to sit with whoever shipped first sits with whoever's about to ship next.
Sources
- https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
- https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html
- https://www.infoworld.com/article/4203865/openai-drops-gpt-5-6-luna-and-terra-api-prices-by-up-to-80.html
- https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost
- https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5