AI Model Report

Reviews · OCTOBER 4, 2026

Gemini 4 Argon tops AutomationBench at 51.3%, setting a new ceiling for end-to-end business execution

Google's new frontier model leads Zapier's end-to-end business benchmark by nine points over Claude Opus 5.5 — but access is limited to trusted cyber defenders at launch.

By Karl Strauchman · Senior model reviewer · October 4, 2026

Google announced Gemini 4 Argon on September 30, 2026, and the headline number is 51.3% on AutomationBench, Zapier's end-to-end business-execution test. That's nine points clear of Claude Opus 5.5 at 42.5% and almost ten clear of GPT-6 Astra at 41.4%. On the benchmark that most closely mirrors how AI actually gets used inside a company, Argon has retaken the top spot.

The broader scorecard is unusually wide. Of 18 disclosed benchmark categories, Argon leads outright in 12 and ties in one. It posts 84.2% on GraphWalks against Astra's 71.8% and Opus 5.5's 66.8%; 65.4% on Vals Finance Agent v2 against 58.6% and 53.5%; 77.9% on DeepSWE v1.1 against roughly 74% for both rivals; and a startling 19.6% on Harvey's Legal Agent Benchmark where Astra scores 5.4% and Opus just 3.8%. CWE-bench v1 is a 68% tie with Astra.

Argon doesn't sweep. GPT-6 Astra still leads FrontierSWE v2 (65.5% to 55.0%) and Terminal-Bench Science 0.1 (68.1% to 57.6%), both by 10.5 points. Claude Opus 5.5 holds Terminal-bench 4.0 (66.4% to 57.4%) and PostTrainBench (49.3% to 45.3%). The pattern is consistent with earlier reporting on AutomationBench's view of Opus 5.5 on real business workflows: frontier models are now specializing in visibly different ways.

Pricing is the other shoe. Argon's introductory rate is $2 per million input tokens and $10 per million output, with a 95% cached-input discount. That's roughly a fifth of Astra's listed $10/$50 and half of Opus 5.5's $4/$20, though Argon reverts to $4/$20 after the introductory window. For buyers who already noted that GPT-6.1 Sol landed at one-fifth of Astra's price while outscoring Opus 5.5, the frontier is now a price war fought on business-workload benchmarks.

The capability jumps are real. Argon raises the output-token ceiling to 1 million from the prior 64,000, and Google says Argon agents replaced 32,000 lines of SIMD code in a Rust port of libgav1 to produce a memory-safe video decoder 2.7x faster than the previous version, plus more than 300 TiB of memory optimizations across its data centers (with 500 TiB to 1 PiB of total savings estimated). Safety posture is tight too: a 0.7% attack success rate on Gray Swan's indirect prompt-injection benchmark.

The caveat is distribution. Argon ships to trusted cyber defenders first, not general API customers. For small-business owners buying done-for-you growth work from services like LemonLime, the ceiling has moved. What hasn't moved yet is who gets to use it.

Sources

  • https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
  • https://techcrunch.com/2026/09/30/google-releases-gemini-4-argon-called-its-most-powerful-model-yet/
  • https://www.cnbc.com/2026/10/01/google-gemini-4-arrives-as-wall-street-shifts-to-personal-agents.html
  • https://www.bloomberg.com/news/videos/2026-10-01/google-rolls-out-ai-model-gemini-4-argon-video
  • https://venturebeat.com/technology/google-unveils-gemini-4-argon-retaking-benchmark-lead-over-openai-and-anthropic-but-in-limited-release