The DecoderOriginal · English

Reflection's Beam becomes the most capable open-weight model built outside China

Reflection has released Beam, its first open-weight model. The mixture-of-experts system activates just 23 billion of its 501 billion parameters per token and aims to match GLM 5.2 on coding and reasoning while using…

Jonathan Kemper
Image source · The Decoder

Reflection's Beam becomes the most capable open-weight model built outside China

Jonathan Kemper Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Oct 6, 2026 Image description Reflection

Key Points

  • AI startup Reflection is releasing Beam, its first open-weight model built for coding and reasoning, with a focus on compute efficiency over raw performance.
  • Beam activates just 23 billion of its 501 billion parameters per token and matches GLM 5.2 on key benchmarks while using three to four times less compute, according to Reflection.
  • The model was trained with reinforcement learning on 10,500 Nvidia GPUs over four weeks. It will ship "later this month" under the Apache 2.0 license.

AI company Reflection has announced Beam, its first freely available model. Instead of chasing peak performance, Beam aims to deliver strong results with minimal compute, putting it in direct competition with Chinese open-weight models from Deepseek and Qwen.

Reflection built Beam for coding, logical reasoning, and agentic tasks. The mixture-of-experts model activates 23 billion of its 501 billion total parameters per token, keeping compute costs relatively low.

On demanding reasoning tasks, Beam matches GLM 5.2 while using three to four times less compute, according to Reflection. The company is going after businesses that use AI for coding and automated workflows but need to keep operating costs in check.

On coding and agent benchmarks, Beam also comes close to the much bigger Qwen3.8-Max. Stronger open models like Kimi K3 still beat it on raw performance, the company says. Reflection says it's already training a successor that aims to close the remaining gap to top-performing open models.Ad

Three scatter plots for DeepSWE v1.1, HLE, and Terminal Bench 2.1 show benchmark scores for Beam, GLM-5.2, Qwen 3.8, Inkling, Nemotron 3 Ultra, and Muse Glimmer plotted against estimated compute cost in PFLOP per attempt.
Across all three benchmarks, Beam matches comparable scores while using far less compute than GLM-5.2 and Qwen 3.8, according to Reflection. | Image: Reflection
Agentic Coding Benchmarks Beam Inkling Nemotron 3 Ultra GLM 5.2 GLM 5.3 Kimi K3 Qwen 3.8 Max DeepSeek V4.1 Flash
DeepSWE v1.1 44.4 NR NR 44.0 61.0 68.0 51.0 74.2
SWE Bench Pro v2-Hard 77.2 56.9 NR NR 84.3 88.2 NR NR
SWE Bench Pro v1 65.5 54.3 46.4 62.1 NR NR 67.7 NR
Terminal Bench v2.1 80.1 63.8 56.4 81.0 88.2 88.3 86.6 90.6
SWE Atlas Codebase QnA 34.6 NR NR NR 61.0 68.0 NR NR
SWEBench Multilingual 78.0 NR 67.7 NR NR NR NR NR
SWEBench Verified 80.9 77.6 70.7 NR NR NR NR NR

Massive reinforcement learning run powered Beam's abilities

Beam gets its capabilities from a combination of standard large-scale pretraining and a particularly compute-heavy reinforcement learning phase. Reflection says it ran 10,500 Nvidia GB300 GPUs for over four weeks during that RL phase, calling it one of the largest training runs any open lab has done. Performance kept improving through the end of the run without hitting a ceiling.

Three line charts show Beam's scores on DeepSWE, HLE, and Terminal Bench 2.1 climbing steadily as the number of reinforcement learning rollouts increases to over 80 million.
Beam's benchmark scores continued to climb throughout reinforcement learning, with no signs of plateauing even at over 80 million rollouts. | Image: Reflection

Users can also control how thoroughly Beam reasons through a problem. A tunable parameter lets you choose whether the model answers quickly or takes more time to think on harder tasks, trading off compute cost against output quality.

Three scatter plots for DeepSWE v1.1, HLE, and Terminal Bench 2.1 show benchmark scores for Beam and competing models plotted against the average number of generated tokens.
Measured by the number of generated tokens, Beam also works more efficiently than GLM-5.2 and Qwen 3.8 at similar performance levels. | Image: Reflection

Reflection observed what it calls "emergent capabilities" during training. While running an RL mix of reasoning, software engineering, and terminal tasks, the company noticed Beam getting better at web browsing even though no browsing tasks were part of that training mix. With web access, the model independently learned to query other language models and pull documents from external services.

Reflection shared several demos, including a live-updating New York City subway map, a small 3D game, and a notebook for fine-tuning another AI model. Beam is text-only, but Reflection says it can process content from other media formats as long as they're represented as text.Ad

A separate model handles safety and alignment

For safety and alignment, Reflection trained a second model and merged it with Beam. The guidelines range from hard rules the model must never break to quality standards like factual accuracy and admitting uncertainty, along with a direct, thorough, and proactive response style. The company plans to publish its safety test results in a technical report and open-source the evaluation methods it developed.

A technical report, developer documentation, and model weights under the open Apache 2.0 license are set to ship "later this month," according to the company. Beam is still going through final safety testing, and an early version is available to select users for now.

Former Deepmind researchers backed by $2 billion

Reflection was founded in 2024 by former Google Deepmind researchers Misha Laskin and Ioannis Antonoglou. Laskin led reward modeling for Gemini, and Antonoglou helped build AlphaGo. The startup launched in March 2025 with $130 million in seed funding and the goal of building superintelligence through autonomous coding. The vision was that language models could learn to act as independently as AlphaGo plays Go, using reinforcement learning on computers.

In summer 2025, the company released Asimov, an agent for analyzing large codebases. Reflection then raised $2 billion at an $8 billion valuation in October 2025, with Nvidia among the investors, and has since positioned itself as a Western counterpart to Deepseek and Qwen. The company releases its model weights but keeps training data and pipelines proprietary. More recently, Reflection signed billion-dollar compute deals with SpaceX and cloud provider Nebius.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Source: Reflection
Original source

The Decoder

Content notes

Original publication and rights belong to the source.