Reflection AI has launched Beam, a new open-weight model designed for coding and agentic workloads. It features 501 billion total parameters but activates only 23 billion per token. The company claims this architecture uses three to four times less inference compute than larger rivals while maintaining performance on reasoning benchmarks.
In this article
Deployment is not yet available for self-hosting. The model is currently undergoing final red-teaming. Early access is restricted to a waitlist on the Reflection platform.
What is Reflection Beam
Beam is a general agent model trained from scratch by Reflection AI. It targets enterprise coding and agentic workloads. The team positions the model as advancing the Western open-weight frontier. Kimi K3 currently leads on raw capability, so Beam focuses on efficiency at inference time.
Users control a reasoning effort parameter. Lower settings produce short answers. Higher settings allow longer reasoning chains on difficult tasks. Teams can match effort to task difficulty and compute budget.
Pretraining details
Beam was pretrained on 23.8 trillion tokens from the web, public sources and proprietary licensed datasets. The research team states that curation removed about 95% of raw internet tokens. It also kept roughly 1.8 trillion high-quality tokens that conventional filters would have dropped.
The architecture interleaves local and global attention with fine-grained routed experts. Load balancing builds on auxiliary-loss-free balancing from DeepSeek-V3, adding cosine decay of expert-bias updates. The busiest expert reached just 1.04x average load at the end of pretraining. Across all 52 layers, residual norms stayed bounded using depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation.
Pretraining finished in under 4 weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Goodput reached 92.3% near the end, with 9 semi-automatic rewinds. Midtraining extended effective context to 1M tokens.
High-Compute Reinforcement Learning
RL is Beam’s central scaling axis. The run used 10.5K NVIDIA GB300 GPUs for 4 weeks and generated over 100 million rollouts. Maximum rollout context was 256K tokens. Training and grading consumed about 1.3 billion sandboxes across nearly 1 million coding, agentic and STEM environments.
Reflection trained with fully asynchronous policy gradients. Every token is tagged with the policy version that produced it. New algorithms kept learning stable even at one-day staleness, 107 weight versions behind the current policy. The team reports no plateau as RL compute increased.
Infrastructure numbers are notable. The system sustained 110K concurrent rollouts on average. New weights reached the inference fleet in a median of about 12 seconds. 71 inference incidents were handled without stopping training.
A controllable length penalty taught Beam to solve tasks with fewer tokens. Browsing skills also improved without browsing tasks in the RL mix, which suggests transfer across agentic domains.
Safety and Alignment
Reflection trained a separate safety and alignment teacher from the pretrained checkpoint. It merged that teacher with the RL teacher using multi-teacher on-policy distillation. Safety training used deliberative alignment. Safety evaluation results will appear in the technical report.
Benchmarks (Reflection-Reported)
On SWE-bench Verified, Beam scores 80.9 versus 70.7 for Nemotron 3 Ultra. On Terminal Bench v2.1, Beam scores 80.1, close to GLM 5.2 at 81.0. DeepSeek V4.1 Flash (90.6) and Kimi K3 (88.3) lead there. These numbers come from Reflection’s table, which sources rival scores from Artificial Analysis and DataCurve.
Beam vs Closest Open-Weight Competitors
| Feature | Reflection Beam | GLM-5.2 | Nemotron 3 Ultra | DeepSeek V4.1 Flash | Kimi K3 |
|---|---|---|---|---|---|
| Developer | Reflection AI (US) | Z.ai (China) | NVIDIA (US) | DeepSeek (China) | Moonshot AI (China) |
| Total params | 501B | ~753B | 550B | 552B backbone + 196B Engram | 2.8T |
| Active params | 23B | ~40B | 55B | 8B prefill / 16B decode | 104B |
| Context | 1M (effective) | 1M | Up to 1M | 1M | 1M |
| Input | Text | Text | Text | Text + image | Text + image |
| License | Apache 2.0 (planned) | MIT | OpenMDW-1.1 | MIT | Kimi K3 License |
| Weights | Later in Oct 2026 | Available | Available | Available | Available |
| Terminal Bench v2.1* | 80.1 | 81.0 | 56.4 | 90.6 | 88.3 |
| Source | Reflection | Hugging Face | NVIDIA | Hugging Face | Hugging Face |
*Scores as published in Reflection’s Beam announcement. Specs verified October 5, 2026.
Beam is the smallest model here by total parameters. Its 23B active count sits below GLM-5.2, Nemotron 3 Ultra and Kimi K3. Apache 2.0 and MIT are standard permissive licenses. Kimi K3’s custom license adds attribution requirements for very large products.
What it means
Developers gain a model with high parameter counts that does not require massive compute to run. The active parameter count allows teams to match effort to task difficulty. Apache 2.0 licensing is planned for later in October 2026.



