AI Model Report

Open Source · JULY 20, 2026

Moonshot ships Kimi K3 at 2.8T parameters — the largest open-weight model, weights due July 27

Beijing-based Moonshot released a sparse-MoE frontier system with a 1M-token context and Kimi Delta Attention, claiming a #2–3 overall finish behind Claude Fable 5 and GPT-5.6 Sol on the company's own eval suite.

By Lars Iverson · Open source & model weights · July 20, 2026

Moonshot announced Kimi K3 on Friday, a 2.8-trillion-parameter sparse mixture-of-experts model that Tom's Hardware calls the largest open-weight system ever released. The full checkpoint ships July 27. Everything between now and then is Moonshot-reported.

The architecture is the pitch. K3 activates 16 of 896 experts per token, a sparsity ratio around 1.8%, and pairs a 1-million-token context window with a new attention mechanism the company calls Kimi Delta Attention, alongside Attention Residuals "designed to improve how information flows across sequence length and model depth." The tech blog, titled "Open Frontier Intelligence," also discloses Stable LatentMoE for expert coordination, Quantile Balancing at the router, plus Per-Head Muon, Sigmoid Tanh Unit, and Gated MLA. Moonshot claims "an approximate 2.5x improvement in overall scaling efficiency compared to Kimi K2" and calls K3 "the world's first open 3T-class model."

On Moonshot's own eval suite (current as of July 16), K3 finishes second or third overall behind Claude Fable 5 and GPT-5.6 Sol, and ahead of Claude Opus 4.8 and GPT-5.5. There's a caveat buried in the methodology: "Claude Fable 5 hit fallbacks on 35% of the tasks in our evaluation, which may have negatively impacted its measured performance." Moonshot ran K3 through its own KimiCode harness, Anthropic's models through Claude Code, and OpenAI's through Codex. Read that as you will.

One third-party number does exist. LMArena's Frontend Code evaluation ranked K3 first at 1,679 points in blind developer testing, ahead of Fable 5.

The pricing tells its own story. K3 lists at $0.30 per million cache-hit input tokens, $3 cache-miss, and $15 output. K2's input rate a year ago was $0.60 per million. K3's output pricing is now five times what K2 charged just to read a token, which is what frontier positioning looks like on a rate card.

Markets read the announcement immediately. Z.ai fell 28% Friday, MiniMax dropped 16%, and Alibaba, a Moonshot backer alongside Tencent, slipped 4%. The Financial Times reported Moonshot is raising fresh capital at $31.5 billion, up from $20 billion in May. "K3 raises the capability ceiling for China AI models, shifting the burden of proof to other independent AI labs," analyst Liu told CNBC.

The unresolved backdrop: in February, Anthropic accused Moonshot of using 3.4 million Claude exchanges to train its models via distillation. Moonshot's disclosures say K3 was trained on "export-grade Nvidia silicon and an unnamed alternative GPU vendor," a phrasing that acknowledges the export-control regime without quite explaining how the compute was assembled.

Open weights on July 27 will collapse most of these questions into something testable. Until then, the largest open-weight model ever announced is also the least independently verified.

Sources