Open Source · OCTOBER 6, 2026
Reflection AI's Beam ships as a 501B open-weight MoE with a 23B active footprint
The Brooklyn startup's first model activates 23 billion of 501 billion parameters per token and claims GLM-5.2-level reasoning at 3–4× lower inference compute. Weights, technical report, and serving stack are promised later this month under Apache 2.0.
Reflection AI announced Beam on October 5, a 501-billion-parameter mixture-of-experts model that activates 23 billion parameters per token and will ship under Apache 2.0 later this month. The Brooklyn startup, founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, is positioning Beam as the first US-origin open-weight model that can trade blows with the top Chinese releases on coding and agentic workloads.
That positioning is the real story. For most of 2026 the open-weight frontier has been a Chinese show, from Z.ai's MIT-licensed GLM-5.3 Flash to DeepSeek's V4.1 Flash release. SiliconAngle calls Beam the first US-origin open-weight model to claim parity with those on coding and agentic benchmarks. TechCrunch's more direct comparison is Inkling, the model Mira Murati's Thinking Machines Lab released in July; Reflection's benchmarks show Beam outscoring it on four coding tests where both report results, though Inkling is multimodal and Beam is text-only.
The headline efficiency claim is 3–4× less inference compute than GLM-5.2, which TechCrunch pegs at roughly 744 billion total and 40 billion active parameters. Superpower Daily notes the comparison counts active parameters and generated tokens while excluding prompt prefill, attention operations, and serving overhead, so it's an estimate rather than measured deployment cost. Reflection draws competing models' evaluations from Artificial Analysis and DataCurve.
The benchmark card is uneven. Beam posts 80.9 on SWEBench Verified, 77.2 on SWE Bench Pro v2-Hard, 97.8 on AIME 2026, and 90.5 on GPQA Diamond. On Terminal-Bench v2.1 it scores 80.1 against Kimi K3's 88.3, and on DeepSWE v1.1 it lands at 44.4 versus Kimi K3's 68.0. Reasoning-heavy and standard SWE: strong. Pure terminal and agentic SWE harnesses: still behind.
The infrastructure disclosure is where Beam's pitch starts to feel like a research artifact as much as a product. Pretraining on 23.8 trillion tokens ran under four weeks across 6,144 NVIDIA GB300 NVL72 GPUs, with 9 semi-automatic rewinds and 92.3% goodput. The RL campaign ran four weeks on 10,500 GB300s, generating more than 100 million rollouts at up to 256,000 tokens of context, across roughly 1.3 billion sandboxes. Unite.AI reports an average of 110,000 concurrent rollouts and up to 170,000 concurrent sandboxes spread across more than 20 clusters, two clouds, and four regions. The context window is one million tokens, extended during midtraining.
TechCrunch notes none of the performance claims have been independently verified. Weights, technical report, model card, and the full stack for running, evaluating, and fine-tuning are scheduled for later in October. Until then, Beam is a well-documented promise, exactly the kind of promise the US open-weight ecosystem has been waiting to make.
Sources
- https://techcrunch.com/2026/10/05/reflection-debuts-beam-a-open-weight-ai-model-to-rival-chinese-models-at-lower-compute-cost/
- https://reflection.ai/blog/introducing-beam
- https://siliconangle.com/2026/10/05/reflection-ai-debuts-open-source-beam-model-with-501b-parameters/
- https://www.unite.ai/reflection-ai-unveils-beam-a-501b-parameter-open-weight-model/
- https://superpowerdaily.com/posts/reflection-introduces-beam-with-early-access-ahead-of-october-weight-release