AI Model Report

Open Source · JULY 20, 2026

Kimi K3 lands at 2.8T parameters, takes Frontend Code Arena, jams Moonshot's own capacity

Moonshot's July 17 open-weight release outscored every model except Claude Fable 5 and GPT-5.6 on its own suite, took the Arena Frontend Code top slot at 1,679 points, and forced the company to pause new subscriptions inside 72 hours.

By Lars Iverson · Open source & model weights · July 20, 2026

Moonshot AI released Kimi K3 on July 17, a 2.8-trillion-parameter sparse mixture-of-experts model that took first place on the Arena Frontend Code leaderboard at 1,679 points, outscored every published model except Claude Fable 5 and GPT-5.6 Sol on Moonshot's own evaluation suite, and inside 72 hours forced the company to pause new Kimi subscriptions because inbound requests in a 48-hour window "sharply exceeded forecasts and were approaching the limits of existing clusters." The full weights are scheduled to drop on July 27.

The market read it as a threshold event. Z.ai closed down 28% on Friday, MiniMax fell 16%, Alibaba dropped 4%. Alibaba's own Qwen3.8-Max-Preview, at 2.4 trillion parameters, was suddenly no longer the frontier Chinese model, and the reweighting was violent.

The architecture is where the interesting reading is. K3 activates 16 of 896 experts per token, roughly 1.8% of the pool, and pairs that sparsity with a 1-million-token context window, a new hybrid linear attention scheme called Kimi Delta Attention, and an inter-layer mechanism Moonshot calls Attention Residuals. The company claims a 2.5x scaling-efficiency gain over Kimi K2. It also shipped its own Triton-like compiler, MiniTriton, and published kernel benchmarks against the Nvidia H200, the export-controlled L20, and an unnamed "GPGPU from an alternative vendor." The subtext isn't subtle: this is a model built to run on whatever silicon Beijing can actually get.

Bloomberg's follow-up on July 20 argued the story is less about compute than memory, pointing at SK Hynix and the HBM supply chain rather than Nvidia. That reframing matters because it locates K3's real bottleneck upstream of the model itself.

Pricing tells a second story. K3's API runs $0.30 per million cache-hit input tokens, $3 uncached, and $15 per million output tokens. K2 launched at $0.60 per million input. Uncached K3 input is five times the K2 price. Moonshot has stopped pretending to compete on cost, which is what companies do when they think they're selling a frontier product.

The context around all this is a $2 billion-plus raise in May, reportedly with Goldman Sachs and CICC circling an IPO. Pausing subscriptions during the loudest week of the year isn't a capacity failure so much as a scarcity signal aimed at the book. DeepSeek R1's January 2025 moment produced a similar reflex in US equities and a similar rush to reframe Chinese labs as legitimate frontier actors; eighteen months on, K3 is the same argument made with a larger parameter count and a live IPO narrative attached.

What's genuinely new is the release posture. The largest open-weight model ever shipped is coming out of Beijing on July 27, priced like a closed one, benchmarked mostly against itself.

Sources

  • https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3
  • https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals
  • https://www.bloomberg.com/news/articles/2026-07-20/moonshot-s-kimi-k3-may-be-more-about-memory-than-compute
  • https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html
  • https://finance.yahoo.com/technology/ai/articles/chinas-moonshot-pauses-kimi-subscriptions-080317250.html