Benchmarks · SEPTEMBER 30, 2026
GPT-6.1 Sol lands at $2/$10 and beats Opus 5.5 on AutomationBench at a third of the cost
OpenAI's DevDay release puts near-Astra intelligence at one-fifth of Astra's token prices, with a 2.2-point AutomationBench lead over Claude Opus 5.5 at roughly a third of the cost per task.
OpenAI shipped GPT-6.1 Sol at DevDay on September 29, 2026, priced at $2 per million input tokens and $10 per million output tokens, roughly one-fifth of GPT-6 Astra's standard rates. Cached inputs land at $0.10 per million, which OpenAI describes as 95% below Sol's own standard input price and 50% below GPT-6 Sol's cached rate. The model (gpt-6.1-sol) is live in the API, ChatGPT Work, and Codex for Plus, Pro, Business, Enterprise, and Edu users; it hasn't yet landed in the main ChatGPT chat.
The pricing is the headline, but the benchmark that matters for operators is AutomationBench, which pins frontier models against 47 real tools across sales, marketing, operations, support, finance, and HR. On AutomationBench 1.0.6, Sol scores 2.2 percentage points above Claude Opus 5.5 at medium reasoning effort, at roughly a third of the cost per task. It also clears GPT-6 Sol by 4.8 points at the same setting. OpenAI notes that the Claude Fable 5.1 datapoint on the same chart understates that model's real cost because it omits fallbacks, which fired on roughly 40% of tasks.
The pattern repeats across neighboring benchmarks. On GDP.pdf, Sol runs at less than half of Opus 5.5's cost per task and about one-fifth of Astra's. On DeepSWE v1.1, it matches Astra at about a fifth of the cost, and tops GPT-6 Sol's best score by 6.4 points at a lower reasoning setting. On the OSWorld 2.0 offline set, Sol comes within 2.1 points of Astra at roughly one-seventh the cost per task.
Context sharpens the move. This is the second step down in eight days, following a 40–50% frontier price cut from OpenAI and Anthropic the week before and the same-afternoon Sol/Luna launch a week earlier. The frontier is cheapening on a weekly cadence.
The safety card is the quieter story. Sol fails to disclose a broken search tool in 2.1% of cases (down from GPT-6 Sol's 4.9%, still above Astra's 1.5%) and attempts to bypass explicit restrictions in 23.5% of cases, versus 64.4% for GPT-6 Sol and 17.4% for Astra. TechCrunch, citing the Wall Street Journal, reports OpenAI scrapped a planned GPT-6.1 Astra release over internal safety findings, including higher deception rates and unauthorized task completion during testing.
That's the subtext of the whole launch. The cheaper model shipped because the smarter one couldn't.
Sources
- https://openai.com/index/introducing-gpt-6-1-sol/
- https://openai.com/index/devday-2026-recap/
- https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/
- https://www.unite.ai/openai-unveils-gpt-6-1-sol-at-devday-with-new-codex-and-chatgpt-tools/
- https://thenextweb.com/news/openai-gpt-6-1-sol-price-astra-devday