AI Model Report

Open Source · JULY 21, 2026

Thinking Machines ships Inkling: a 975B open-weights MoE that isn't trying to win

Mira Murati's first production model debuts at 41 on the Artificial Analysis Intelligence Index — third-highest among open weights, ahead of Nemotron 3 Ultra, and explicitly positioned as a fine-tuning base rather than a leaderboard entry.

By Lars Iverson · Open source & model weights · July 21, 2026

Thinking Machines Lab released Inkling on Wednesday, a 975B-parameter multimodal Mixture-of-Experts model with 41B active parameters and a 1M-token context window, published under an Apache-style open-weights license roughly nine months after the lab's first research previews. It landed at 41 on the Artificial Analysis Intelligence Index, third-highest among open weights and comfortably ahead of Nvidia's Nemotron 3 Ultra at 38, Gemma 4 31B at 29, and gpt-oss-120b at 24.

The release copy does something U.S. labs almost never do. It concedes, in print, that Inkling is "not the strongest overall model available today, open or closed." Mira Murati's stealth-era shop isn't shipping a leaderboard entry. It's shipping a fine-tuning base.

That framing matters because the leaderboard story is more complicated than the headline number. On GDPval-AA v2, an agentic Elo eval, Inkling scores 1238 against 1190 for Kimi K2.6 and 1189 for DeepSeek v4 Flash max. On τ³-Banking it hits 24%, edging DeepSeek's 23% and Kimi's 21%. And it does the work efficiently: 25K average output tokens per Intelligence Index task versus 43K for GLM-5.2 max, 38K for Kimi K2.6, and 37K for DeepSeek v4 Pro max. Thinking Machines separately claims Inkling matches Nemotron 3 Ultra's coding performance using a third of the tokens.

The architecture, per the Hugging Face model card, is decoder-only with relative attention rather than RoPE, a fourth projection producing a per-token, per-head relative feature tweaked with key-query distance, and global/sliding-window layers alternating in a 5:1 pattern. The router jointly scores routed and shared experts, with top-k selection over six routed experts plus two shared experts always active. Vision inputs pass through a hierarchical MLP patchifier; audio is encoded as a discretized mel spectrogram, chunked into 100 ms segments and classified into mel bins. Training ran on 45 trillion tokens of text, images, audio and video, using Nvidia GB300 NVL72 systems under the gigawatt Vera Rubin partnership Thinking Machines struck with Nvidia in March. Day-zero support ships for transformers, SGLang, and llama.cpp, in BF16 and NVFP4 variants with speculative MTP layers for faster inference.

The commercial thesis sits alongside Tinker, Thinking Machines' fine-tuning platform, and lands into a growing surface of downstream integrators like Glean, Dust, and LemonLime, whose no-code posture makes an open base with credible efficiency numbers unusually attractive. The self-reported Bridgewater Associates collaboration is the pitch in miniature: 84.7% on financial reasoning tests at roughly one-fourteenth the running cost of a proprietary equivalent.

There's a longer arc here that echoes Meta's original Llama 2 gambit in 2023, when releasing weights was a way to shape the developer stack rather than win a benchmark. Murati is running the same play with better numbers and a more honest headline. The open-weights race in the U.S. spent 2025 being ceded to Chinese labs. Inkling is the first credible domestic answer, and it arrives with its ambitions deliberately downgraded.

Sources