AI Model Report

Reviews · JULY 19, 2026

Moonshot's Kimi K3 lands at 2.8T parameters, claims second only to Fable 5 and GPT-5.6 Sol

Beijing's Moonshot AI released Kimi K3 on July 16 with a 1M-token context, Kimi Delta Attention, and $3/$15 input-output pricing. Vendor benchmarks put it behind only Claude Fable 5 and GPT-5.6 Sol; full weights are promised by July 27.

By Karl Strauchman · Senior model reviewer · July 19, 2026

Moonshot AI released Kimi K3 on July 16, a 2.8-trillion-parameter mixture-of-experts model the company positions behind only Claude Fable 5 and GPT-5.6 Sol on its own evaluation suite. Full weights are scheduled to land on July 27, which means the entire narrative currently in market rests on numbers Moonshot itself is publishing.

The architecture is the interesting part. K3 activates 16 of 896 experts per token, routing roughly 1.8% of the pool per forward pass, and pairs that sparsity with Kimi Delta Attention (KDA), a hybrid linear attention scheme, plus Attention Residuals and the Stable LatentMoE training framework. Moonshot claims about 2.5x the scaling efficiency of Kimi K2 and says quantization-aware training in MXFP4 weights and MXFP8 activations began at the SFT stage, chosen for broad hardware compatibility. Context window is one million tokens. The model is live across Kimi.com, Kimi Work, Kimi Code, and the Kimi API.

Pricing tells its own story. K3 lists at $0.30 per million cache-hit input tokens, $3 per million cache-miss inputs, and $15 per million output tokens. K2 launched at $0.60 per million input tokens, so uncached input is now roughly 5x more expensive. Against Fable 5's $50 per million output, K3 still comes in at less than a third. Moonshot is charging what a frontier lab charges, then undercutting the actual frontier by an order of magnitude on egress.

The independent data point most worth watching: K3 took first in Arena's blind Frontend Code evaluation at 1,679 points, ahead of Fable 5. Moonshot's internal ranking across FrontierSWE, PostTrain Bench, and MLS Bench Lite, run through KimiCode, Claude Code, and Codex harnesses, places K3 third overall behind Fable 5 and GPT-5.6 Sol, with Claude Opus 4.8, GPT-5.5, and GLM-5.2 trailing.

Markets read the release as a sequel. TSMC fell 7% Friday despite quarterly operating profit jumping 77%, SoftBank dropped 9%, and Nvidia shed 1.2%, a pattern Fortune framed as a second DeepSeek moment. Bank of America analysts led by Alex Liu argued in a note cited by CNBC that "pre-training scaling, paired with architectural innovation, can still deliver step-change gains for flagship Chinese models."

The overhang is trust. In February, Anthropic accused Moonshot of training on 3.4 million Claude exchanges via distillation, and Tom's Hardware bluntly reminded readers that "Every published K3 number is a claim made by Moonshot or drawn from API access and can't be verified until the weights are made public on July 27."

Nine days is a short wait for a large answer.

Sources

  • https://www.kimi.com/blog/kimi-k3
  • https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3
  • https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html
  • https://fortune.com/2026/07/17/china-moonshot-kimi-k3-markets-china-ai/
  • https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals