Open Source · OCTOBER 6, 2026
Reflection AI unveils Beam, a 501B open-weight MoE claiming 3–4× less inference compute than GLM-5.2
The Nvidia-backed Brooklyn lab announced its first model on October 5, promised Apache 2.0 weights later this month, and positioned a 501B/23B-active MoE pretrained on 23.8 trillion tokens against DeepSeek, Qwen, and Z.ai.
Reflection AI announced Beam on October 5: a 501-billion-parameter sparse mixture-of-experts with 23 billion active parameters, pretrained on 23.8 trillion tokens, aimed squarely at the coding, reasoning, and agentic workloads where Chinese open-weight labs have spent the last year setting the pace. The Brooklyn lab is positioning Beam less as a frontier flagship than as an efficiency play, and the framing is itself the story.
The headline claim is that Beam needs 3–4× less inference compute than Z.ai's GLM-5.2 at comparable reasoning quality, with larger gains against the Qwen 3.8-Max family north of two trillion total parameters. GLM-5.2 is reportedly around 744B total and 40B active per TechCrunch and Reuters, so the architectural pitch has a visible referent. Reflection also publishes a benchmark table: 80.9 on SWEBench Verified, 80.1 on Terminal Bench v2.1, 77.2 on SWE Bench Pro v2-Hard, 90.5 on GPQA Diamond, and 36.2 on Humanity's Last Exam without tools.
All of it's first-party. TechCrunch notes the performance claims haven't been independently verified, and Reflection itself concedes its FLOPs figures approximate generation compute as twice the active parameter count times mean generated tokens per attempt, excluding prompt prefill and attention, and describes them as "an approximate compute comparison rather than measured inference cost." Beam's own table shows it trailing Kimi K3 on raw capability, with Z.ai's GLM-5.3 and DeepSeek V4.1 Flash ahead on some agentic-coding benchmarks.
The Apache 2.0 weights, a technical report, model card, and the full stack for running, evaluating, and fine-tuning are promised later this month. Until then, operators can join a waitlist. The gap between announcement and drop is the entire ballgame; it's also where credibility gets settled.
What's not in doubt is the capital stack. Reflection has raised roughly $4.7 billion from Nvidia, Sequoia, and Lightspeed at a $25 billion pre-money valuation, and signed more than $7 billion in compute deals with SpaceX and Nebius running through 2029. The RL run that produced Beam used 10,500 Nvidia GB300 GPUs over four weeks and more than 100 million rollouts. Fortune reports a larger successor is already training.
Read structurally, Beam is the first serious Western answer to the thesis that open weights, not closed APIs, are where cost-performance competition now happens, a thesis Chinese labs built by shipping. Reflection is attempting to buy it back with Nvidia silicon and a permissive license. The weights will arrive, or they won't, and that's the only test that matters.