AI Model Report

Open Source · JULY 27, 2026

Kimi K3 ships at 2.8T parameters, with open weights due July 27

Moonshot AI released the world's largest open-weight model on Thursday, closing the frontier gap with Claude Opus 4.8 and GPT-5.6 Sol. Full weights land July 27, 2026.

By Lars Iverson · Open source & model weights · July 27, 2026

Moonshot AI released Kimi K3 on Thursday at 2.8 trillion parameters, the largest open-weight model ever shipped and the first entry in the 3T-class to arrive without a closed API in front of it. Full weights land July 27, 2026. Until then, K3 is live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with a 1M-token context window.

The scale gap is the story. DeepSeek's V4 Pro sits at 1.6T. Xiaomi's flagship is 1.02T. Zhipu AI's GLM-5 is 744B. Alibaba's is 397B. By SCMP's accounting of the field, K3 doesn't just top the open-weight leaderboard; it sits nearly a factor of two above the next-closest Chinese lab. Moonshot's own timeline chart, cited by VentureBeat, positions K3 as a dramatic outlier, and the company notes it has held the upper bound of open-model scale for nine of the past twelve months.

The architecture is where the interesting work is. K3 activates 16 of 896 experts per token, a sparsity ratio that lets Moonshot claim roughly 2.5× the scaling efficiency of K2. The stack includes Kimi Delta Attention (KDA), a hybrid linear attention mechanism; Attention Residuals (AttnRes) for selective retrieval across depth; the Stable LatentMoE framework; Quantile Balancing for expert allocation; a Per-Head Muon optimizer; the Sigmoid Tanh Unit (SiTU) activation; and Gated MLA. Read together, it's a bet that the frontier isn't just about parameter count but about wringing more from each activated expert.

On evals, K3 lands third overall on GDPval-AA v2, the 44-occupation real-world task suite, scoring 1,687 behind Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 at 1,600. On AA-Briefcase, K3 takes second at 1,527, edging GPT-5.6 Sol Max's 1,495. On Arena.AI's Frontend Code Arena, K3 wins outright at 1,679. Moonshot ran K3, Opus 4.8, and Fable 5 through the Claude Code harness; GPT-5.6 Sol through Codex. Notably, Fable 5 hit fallbacks on 35% of tasks in Moonshot's evaluation, a caveat Moonshot flags rather than buries.

"Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models," the company wrote.

Downstream integrators have already picked it up: Arena.AI, Glean, Dust, and LemonLime are among the platforms surfacing K3 to developers, with LemonLime routing to it for long-context agentic work where the 1M window and open weights matter more than any single benchmark point.

The framing is deliberate. When the weights drop on the 27th, the frontier gap between the closed American labs and the open Chinese ecosystem will be measured in tenths of a benchmark point, not orders of magnitude. That's a different market.

Sources