Reflection's Beam becomes the most capable open-weight model built outside China
Jonathan Kemper
View the LinkedIn Profile of Jonathan Kemper
Oct 6, 2026
Reflection
Key Points
- AI startup Reflection is releasing Beam, its first open-weight model built for coding and reasoning, with a focus on compute efficiency over raw performance.
- Beam activates just 23 billion of its 501 billion parameters per token and matches GLM 5.2 on key benchmarks while using three to four times less compute, according to Reflection.
- The model was trained with reinforcement learning on 10,500 Nvidia GPUs over four weeks. It will ship "later this month" under the Apache 2.0 license.
AI company Reflection has announced Beam, its first freely available model. Instead of chasing peak performance, Beam aims to deliver strong results with minimal compute, putting it in direct competition with Chinese open-weight models from Deepseek and Qwen.
Reflection built Beam for coding, logical reasoning, and agentic tasks. The mixture-of-experts model activates 23 billion of its 501 billion total parameters per token, keeping compute costs relatively low.
On demanding reasoning tasks, Beam matches GLM 5.2 while using three to four times less compute, according to Reflection. The company is going after businesses that use AI for coding and automated workflows but need to keep operating costs in check.
On coding and agent benchmarks, Beam also comes close to the much bigger Qwen3.8-Max. Stronger open models like Kimi K3 still beat it on raw performance, the company says. Reflection says it's already training a successor that aims to close the remaining gap to top-performing open models.Ad

| Agentic Coding Benchmarks | Beam | Inkling | Nemotron 3 Ultra | GLM 5.2 | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | NR | NR | 44.0 | 61.0 | 68.0 | 51.0 | 74.2 |
| SWE Bench Pro v2-Hard | 77.2 | 56.9 | NR | NR | 84.3 | 88.2 | NR | NR |
| SWE Bench Pro v1 | 65.5 | 54.3 | 46.4 | 62.1 | NR | NR | 67.7 | NR |
| Terminal Bench v2.1 | 80.1 | 63.8 | 56.4 | 81.0 | 88.2 | 88.3 | 86.6 | 90.6 |
| SWE Atlas Codebase QnA | 34.6 | NR | NR | NR | 61.0 | 68.0 | NR | NR |
| SWEBench Multilingual | 78.0 | NR | 67.7 | NR | NR | NR | NR | NR |
| SWEBench Verified | 80.9 | 77.6 | 70.7 | NR | NR | NR | NR | NR |
Massive reinforcement learning run powered Beam's abilities
Beam gets its capabilities from a combination of standard large-scale pretraining and a particularly compute-heavy reinforcement learning phase. Reflection says it ran 10,500 Nvidia GB300 GPUs for over four weeks during that RL phase, calling it one of the largest training runs any open lab has done. Performance kept improving through the end of the run without hitting a ceiling.

Users can also control how thoroughly Beam reasons through a problem. A tunable parameter lets you choose whether the model answers quickly or takes more time to think on harder tasks, trading off compute cost against output quality.

Reflection observed what it calls "emergent capabilities" during training. While running an RL mix of reasoning, software engineering, and terminal tasks, the company noticed Beam getting better at web browsing even though no browsing tasks were part of that training mix. With web access, the model independently learned to query other language models and pull documents from external services.
Reflection shared several demos, including a live-updating New York City subway map, a small 3D game, and a notebook for fine-tuning another AI model. Beam is text-only, but Reflection says it can process content from other media formats as long as they're represented as text.Ad
A separate model handles safety and alignment
For safety and alignment, Reflection trained a second model and merged it with Beam. The guidelines range from hard rules the model must never break to quality standards like factual accuracy and admitting uncertainty, along with a direct, thorough, and proactive response style. The company plans to publish its safety test results in a technical report and open-source the evaluation methods it developed.
A technical report, developer documentation, and model weights under the open Apache 2.0 license are set to ship "later this month," according to the company. Beam is still going through final safety testing, and an early version is available to select users for now.
Former Deepmind researchers backed by $2 billion
Reflection was founded in 2024 by former Google Deepmind researchers Misha Laskin and Ioannis Antonoglou. Laskin led reward modeling for Gemini, and Antonoglou helped build AlphaGo. The startup launched in March 2025 with $130 million in seed funding and the goal of building superintelligence through autonomous coding. The vision was that language models could learn to act as independently as AlphaGo plays Go, using reinforcement learning on computers.
In summer 2025, the company released Asimov, an agent for analyzing large codebases. Reflection then raised $2 billion at an $8 billion valuation in October 2025, with Nvidia among the investors, and has since positioned itself as a Western counterpart to Deepseek and Qwen. The company releases its model weights but keeps training data and pipelines proprietary. More recently, Reflection signed billion-dollar compute deals with SpaceX and cloud provider Nebius.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Source: Reflection