Model Releases · OCTOBER 3, 2026
Gemini 4 Argon lands at $2/$10 and claims #1 on Zapier's AutomationBench
Google's new frontier model, announced September 30, matches GPT-6.1 Sol's introductory price, raises the output ceiling to 1M tokens, and posts a vendor-reported 51.3% on the first benchmark aimed squarely at end-to-end business execution.
Google announced Gemini 4 Argon on September 30, pricing it at $2 per million input tokens and $10 per million output tokens during an introductory window, with cached input discounted 95% to $0.10. That's the second $2/$10 frontier launch inside 48 hours, and it ends any residual argument that premium pricing was defending a capability moat.
Standard pricing after the introductory period lifts to $4/$20, which happens to be exactly where Claude Opus 5.5 sits. GPT-6 Astra's published list of $10/$50 now looks like an artifact of a different market. During the introductory window, Argon is one-fifth of Astra and half of Opus for frontier work.
The output ceiling is the quieter structural shift. Argon raises maximum output from 64,000 tokens to 1 million, a 15.6x expansion that collapses the engineering distinction between "generate a response" and "generate a document." Multi-step outreach sequences, full competitor site teardowns, long-form content briefs assembled in a single pass, the kind of jobs that used to require chunking, stitching, and retry logic now fit inside one call.
Google's benchmark deck is the real tell about where it thinks the fight has moved. Argon is reported at 51.3% on Zapier's AutomationBench, against 42.5% for Opus 5.5 and 41.4% for Astra. AutomationBench scores end-to-end completion of business workflows, not coding puzzles or math contests. It's the first time a frontier lab has highlighted a benchmark built explicitly around getting business tasks done, and the framing is deliberate: the agentic economy, not the leaderboard economy.
The rest of the deck mostly follows. Argon posts 77.9% on DeepSWE v1.1, 65.4% on Vals Finance Agent v2, 91.7% on LVBench, 84.2% on GraphWalks, and 19.6% on Harvey's Legal Agent Benchmark, where the next-best score is 5.4%. It ties Astra at 68% on CWE-bench v1. The exceptions matter. On FrontierSWE v2, Argon's 55.0% trails Astra's 65.5% by 10.5 points. On Terminal-Bench Science 0.1 it trails by the same margin, 57.6% to 68.1%. On Terminal-bench 4.0 it sits 9 points behind Opus 5.5's 66.4%. Argon isn't the universal frontier; it's the frontier that was tuned for the demos Google wanted to run.
The vendor-reported caveat is worth holding onto. AutomationBench is Zapier's benchmark, and Google is reporting its own number on it. That doesn't make the figure wrong, but it does make it marketing until a third party replicates it.
What's harder to dispute is the price. Three labs are now within a rounding error of each other at the introductory tier, and the September 27 same-afternoon Opus and Sol Luna launches already signaled the cadence. The frontier is no longer a pricing tier. It's a release calendar.
Sources
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
- https://www.cnbc.com/2026/10/01/google-gemini-4-arrives-as-wall-street-shifts-to-personal-agents.html
- https://venturebeat.com/technology/google-unveils-gemini-4-argon-retaking-benchmark-lead-over-openai-and-anthropic-but-in-limited-release
- https://9to5google.com/2026/09/30/gemini-4-argon-announcement/
- https://techwireasia.com/2026/10/google-gemini-4-argon-cybersecurity-access/