Reviews · SEPTEMBER 13, 2026
Cognition's SWE-2 lands at 50.0% on FrontierCode 1.1 for 64% less, and Fusion arrives in Desktop and CLI
SWE-2, post-trained from Moonshot's 2.8T-parameter Kimi K3 in a single cost-penalized RL run, matches Fable 5.1 within a point at a fraction of the cost. A day later, Cognition shipped its dual-model Fusion harness to Devin Desktop and CLI.
Cognition shipped SWE-2 into Devin Desktop and the Devin CLI on September 10, and expanded its dual-model Fusion harness into the same surfaces a day later. On FrontierCode 1.1 Main, the new coding model scores 50.0%, within a point of Anthropic's Fable 5.1 at 50.9% and trailing OpenAI's GPT-6 Astra at 53.3%. Per task, it runs 64% cheaper than Fable and roughly a quarter the cost of Astra.
That's the headline. The interesting part is how Cognition got there, and what the release says about where the coding-agent market is sorting itself.
SWE-2 is post-trained from Moonshot's Kimi K3, the 2.8-trillion-parameter mixture-of-experts base that landed in July, roughly 3x the parameter count of the Kimi K2.7 that sat under SWE-1.7. Cognition trained all three effort levels (medium, high, max) in a single RL run using a cost-penalized reward of the form R = S − λ_e·C. Appendix B of the company's writeup argues the linear form is the only shape that makes the objective depend solely on average cost and average solve rate. It's a small but pointed piece of research theater: the math justifies the price tag.
Running RL at multi-trillion-parameter scale required four stack changes. A prefill delayer that batches nearby GPU prefill requests bought a 10–20% throughput lift. DSpark and SpecForge, Cognition's speculative decoding system and draft-model trainer, extended accepted sequences by 15%. NVFP4 and FP8 quantization pulled memory down. And the team tripled its RL environment count versus SWE-1.7.
The behavioral shift shows up in steps. SWE-1.7 averaged 127 steps per FrontierCode task; SWE-2 lands at 53, 80, and 98 at medium, high, and max effort. Its median first substantive code edit arrives at step 18, down from step 48. The agent stops thinking out loud and starts editing.
For a five-person B2B engineering team, that arithmetic matters more than the benchmark. Cheaper per-task coding is the same structural story running through DeepSeek's V4.1 Flash and its causal encoder-decoder work: frontier-adjacent quality at commodity prices, distributed through whichever harness the buyer already uses. Fable and Astra still win the top of the leaderboard. But the leaderboard isn't where small teams shop.
The remaining friction is distribution. SWE-2 and Fusion live inside Devin. That's the moat and the ceiling in the same sentence.
Sources
- https://cognition.com/blog
- https://www.marktechpost.com/2026/09/12/cognition-releases-swe-2-a-kimi-k3-post-trained-coding-model-that-matches-fable-5-1-on-frontiercode-at-64-lower-cost/
- https://www.techtimes.com/articles/327315/20260911/cognition-swe-2-beats-frontier-coding-ai-64-lower-cost-using-single-run-rl-training.htm
- https://cryptobriefing.com/cognition-devin-fusion-multi-model-coding-agent/
- https://superpowerdaily.com/posts/cognition-releases-swe-2-for-devin-claiming-lower-cost-coding-performance