Benchmarks · OCTOBER 3, 2026
Gemini 4 Argon lands at $2/$10 with a 51.3% AutomationBench — and no API ID yet
Google's new frontier model tops Zapier's end-to-end business-task benchmark and matches GPT-6.1 Sol on introductory price, but access is gated to Fairwind cyber defenders while the broader API rollout date remains unset.
Google announced Gemini 4 Argon on September 30, posting a 51.3% score on AutomationBench, Zapier's end-to-end business-task benchmark, roughly nine points clear of the nearest frontier rival. Of the 18 benchmarks Google disclosed, VentureBeat counts Argon leading outright on 12 and tying on one. GPT-6 Astra leads three; Claude Opus 5.5 leads two.
The headline price is $2 per million input tokens and $10 per million output tokens, an introductory rate that matches GPT-6.1 Sol and undercuts Claude Opus 5.5's $4/$20 by half on output. Google has already signaled the standard rate will land at $4/$20. Cached inputs get a 95% discount off the input rate, and the output ceiling jumps to 1 million tokens, up from the 64,000 that bounded prior Gemini releases.
For small teams weighing AutomationBench as a signal for sales and marketing automation, that 51.3% matters more than the model's name. AutomationBench measures real workflow execution, which is the category that touches CRM wiring, prospect research pipelines, and content ops. Opus 5.5 scored 42.5%, Astra 41.4%, Claude Fable 5.1 31.4%. The gap is wide enough to notice; the pricing makes it hard to ignore.
The caveat is access. Argon's broader API is gated: current availability runs through the Fairwind Program, which Google opened on September 2 to cyber defenders, and there's no public API ID yet. Developers who want to baseline now should test against Gemini 3.8 Flash, priced at $0.75/$3.75 per million tokens through December 31, 2026, and plan budgets at Argon's standard $4/$20.
Not every benchmark favors Argon. On the Artificial Analysis Intelligence Index, it scores 53, tied with Astra and Fable 5.1 and trailing Opus 5.5 at 58. It also burns roughly 62,000 output tokens per index task, versus Astra's 27,000, which lands Argon at $1.99 per task at introductory pricing and $3.98 at standard, against Astra's $3.26. On Arena's Agent Arena leaderboard as of October 1, Argon sits 8th overall on 3,417 sessions, with a confidence range stretching from 3rd to 16th. Harvey's Legal Agent Benchmark tells a different story again: 19.6% for Argon, 5.4% for Astra, 3.8% for Opus 5.5. On Vals Finance Agent v2, Argon leads at 65.4% over Astra's 58.6% and Opus 5.5's 53.5%.
That shape, dominant on agentic business tasks, middling on raw intelligence indices, verbose per task, mirrors the positioning that GPT-6.1 Sol took when it undercut Astra on price one day earlier. The frontier isn't competing on IQ anymore. It's competing on which model does the work small teams were already paying someone, or something, to do.
Sources
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
- https://www.cnbc.com/2026/10/01/google-gemini-4-arrives-as-wall-street-shifts-to-personal-agents.html
- https://www.bloomberg.com/news/videos/2026-10-01/google-rolls-out-ai-model-gemini-4-argon-video
- https://venturebeat.com/technology/google-unveils-gemini-4-argon-retaking-benchmark-lead-over-openai-and-anthropic-but-in-limited-release
- https://www.rdworldonline.com/googles-overdue-gemini-4-argon-reaches-the-frontier-with-mixed-results-for-science/