Reflection AI has announced its first open-weight model — Beam, a large sparse MoE network that, according to the lab, sets records among open models on coding benchmarks and trails its competitor Kimi K3 only in raw computational power. The weights are promised under the Apache 2.0 license this month, along with a full stack for deployment and fine-tuning, and the main selling point — significantly lower inference compute costs — currently relies on the lab's own counting methodology, which independent researchers are almost certain to challenge.

image

What happened

On October 5, 2026, Reflection AI published a post on its official blog, “Introducing Beam: Reflection's 501B open-weight model,” presenting its first open-weight model. Beam is built as a sparse Mixture-of-Experts: 501 billion parameters in total, of which 23 billion are active simultaneously. Pre-training was conducted on 23.8 trillion tokens, and after mid-training, the effective context reaches 1 million tokens. The developers call the RL stage one of the largest among open labs: over 100 million rollouts in 4 weeks on 10,500 NVIDIA GB300 accelerators, about 1.3 billion sandboxes, and roughly 1 million environments. According to the announcement, fully asynchronous policy gradients remain stable even when the policy weights are 107 versions stale, and one run survived 71 inference incidents. According to the lab's own numbers, Beam shows the best result among open models on SWEBench Verified (80.9) and SWEBench Multilingual (78.0), scores 97.8 on AIME 2026 and 90.5 on GPQA Diamond, but loses to Kimi K3 on SWE Bench Pro v2-Hard — 77.2 versus 88.2.

Context

The open frontier in coding and agentic tasks in recent months has been shaped around Chinese labs: GLM 5.2 and 5.3, Kimi K3, and Qwen 3.8 Max set the bar against which Beam is directly compared. Reflection AI consciously shifts the competition from the plane of raw benchmarks to the plane of inference efficiency, claiming a 3–4-fold reduction in FLOPs at comparable reasoning scores. The approximate estimation formula shown in the blog — roughly 2 × active parameters × generated tokens — does not account for prefill, attention, and serving overhead; with a context of up to 1 million tokens, the prefill share can be significant, so such a count inherently favors MoE architectures with a small number of active parameters. Judging by the published specifications, the architecture itself is a standard frontier configuration without obvious novelty, and the main candidate for reproducible engineering value is the large-scale reinforcement learning infrastructure, not the model itself.

Why this matters for the industry

For the industry, this is Reflection AI's first open frontier release and a claim to be a Western alternative to GLM 5.2/5.3, Kimi K3, and Qwen 3.8 Max in coding and agentic tasks. If the efficiency thesis is confirmed by independent measurements, the unit economics of agentic services will improve, and closed APIs will face price pressure, because the cost of a solved task will be measured in FLOPs and tokens. The second asset is the RL infrastructure: asynchronous policy gradients that tolerate significant staleness of policy weights, combined with sandboxes on the scale of about a billion, form a reproducible engineering base that future open models can reuse; over time, heavy agentic RL will cease to be the privilege of closed labs. At the same time, the FLOPs counting methodology will almost certainly become a subject of debate over fair comparison of MoE model inference costs: there is no comparison on an aligned compute budget in the announcement.

Why this matters for users

For readers working with open models, the main practical fact is the promise to publish the weights under the Apache 2.0 license this month, along with a full stack for deployment, evaluation, and fine-tuning; a 501-billion MoE with 23 billion active parameters belongs to the class of models that enthusiasts with several GPUs can actually run locally. Registration for early access on platform.reflection.ai is already open, where you can try the user parameter reasoning effort, which controls the balance between the length of the model's reasoning and the quality of the answer. The lab also shows a demo: a deterministic 180×90 puzzle with 95.5% coverage and an example of fine-tuning Gemma-4. A reasonable action today is to join the early access queue and prepare your own evaluation methodology on your agentic tasks in advance, so that on release day you can compare Beam by the cost of a solved task, not by the announcement's numbers.

What is still unknown / limitations

There are no weights, technical report, or model card yet — everything is promised for this month, so neither benchmark scores nor efficiency comparisons can be independently verified: we are looking at the lab's own numbers, not third-party runs. The FLOPs counting methodology used in the announcement is biased in favor of Reflection's architecture with a small number of active parameters, and there is no comparison on an aligned compute budget in the material, so talk of a “3–4 times cheaper model” is premature until the weights are released. AIME 2026 and GPQA Diamond are close to saturation and poorly differentiate models, and demonstrative demos like the puzzle and fine-tuning are illustrations, not standardized benchmarks. Independent verification — third-party runs of SWEBench and SWE Bench Pro, real inference cost accounting for prefill and serving, reproduction of RL details — is possible only if the Apache 2.0 weights are released on schedule.

Sources

Author

Look at AI, editorial team