AI Model Report

Reviews · AUGUST 15, 2026

Grok 4.6 ties GPT-5.6 Sol on AA Intelligence Index at $2/$6, but trails on DeepSWE and Terminal-Bench

SpaceXAI's post-training refresh keeps the Grok 4.5 base, adds an xhigh reasoning effort and 500K context, and scores 61 on Artificial Analysis — one point behind Claude Fable 5 and roughly a third the price of comparable frontier rivals.

By Karl Strauchman · Senior model reviewer · August 15, 2026

SpaceXAI shipped Grok 4.6 on August 12, 2026, and the headline number is 61 on the Artificial Analysis Intelligence Index, a nine-benchmark composite that now places Grok tied with OpenAI's GPT-5.6 Sol and exactly one point behind Anthropic's Claude Fable 5. That's the story the release notes want you to read. The pricing tells a different one: $2 per million input tokens and $6 per million output, roughly a third of what comparable frontier rivals charge.

Under the hood, this isn't a new base. SpaceXAI kept Grok 4.5 as the foundation and spent the improvement budget on a longer supplemental training run, curated model-generated reasoning data, higher-quality engineering data, and a reworked optimizer and training recipe. Grok 4.5 itself was used to regenerate SFT trajectories across reasoning-effort levels, agent harnesses, and STEM, software-engineering and knowledge-work domains, with problematic traces filtered by model-based checks. Context expands to 500,000 tokens, text and image in, text-only out, with a February 1, 2026 knowledge cutoff.

The composite flatters the model. As our desk read of the leaderboard puts it: "A tie on a composite is not a tie on every constituent. Grok 4.6's coding-eval deficits are visible in the same table that produces the 61." MarkTechPost notes Grok 4.6 trails on DeepSWE and Terminal-Bench, the two coding-agent benchmarks that most closely map to the actual work enterprise buyers are trying to automate. Across nine additional benchmarks compared against Claude Fable 5, per SiliconAngle, the Grok family topped Anthropic's flagship on three. VentureBeat framed the release accurately: Grok 4.6 vaulted past Kimi K3 to fourth on Artificial Analysis rather than establishing an uncontested lead.

Pricing has structure worth reading carefully. Below a 200K prompt-token threshold, it's $2/$0.50 cached/$6. Above it, the rate doubles to $4/$1 cached/$12. A fast variant charges double the standard rate. MarkTechPost also flags an operational trap: teams that don't set a prompt_cache_key or x-grok-conv-id header on Chat Completions will see requests scatter across servers, break cache hits, and pay the full input rate. That's a real bill, quietly.

Distribution is where the strategy sharpens. SpaceXAI acquired Cursor in June for $60 billion, and Grok 4.6 arrives inside that harness with 2x included usage for the first week through both Cursor and Grok Build. VentureBeat's read is that the Cursor placement matters for enterprise buyers who want a coding model dropped into an existing harness rather than a new API integration, a shortcut around the coding-eval deficit the same benchmarks just documented.

SpaceXAI's qualitative claim that Grok 4.6 self-tests and verifies more on long trajectories is, per MarkTechPost, a vendor observation from internal testing rather than an independently measured result. VentureBeat also flags the Grok brand's documented history of safety incidents as a procurement consideration separate from the technical result. Fourth place, priced to sell, sold through the editor you already use.

Sources