
Reflection AI dropped Beam on October 5 — a 501 billion parameter open-weight model aimed at coding and agentic work. That number is almost designed to mislead you. Beam is a sparse Mixture-of-Experts model, which means only 23 billion parameters activate during inference. The 501B figure describes total weights on disk; your GPU only touches a fraction of them per forward pass. That distinction matters because it is the source of Reflection’s core claim: Beam delivers GLM-5.2-level reasoning at three to four times less inference compute. Whether that claim survives real-world testing remains open — the weights are not out yet.
What Beam Actually Is
Beam was pretrained on 23.8 trillion tokens across 6,144 NVIDIA GB300 GPUs in under four weeks. That is the foundation. The more interesting investment came after: Reflection ran over 100 million reinforcement learning rollouts on 10,500 GPUs using roughly 1.3 billion sandboxed evaluation environments. The company describes it as one of the largest RL training runs by any open lab. The scale shows up in agentic behavior — during training, the model independently discovered it could query external language models and invoke OCR APIs, capabilities not explicitly trained.
The model supports a claimed context window of up to 1 million tokens and includes an adjustable reasoning effort parameter, letting developers dial compute depth from low to max on a per-request basis. The API endpoint follows OpenAI’s chat completions format, which means no new tooling is required for integration.
Benchmarks: Where Beam Wins and Where It Doesn’t
Reflection’s benchmark comparisons deserve scrutiny. On reasoning tasks, Beam posts legitimately strong numbers: 97.8 on AIME 2026, 90.5 on GPQA Diamond, and 80.9 on SWE-bench Verified. The comparisons Reflection highlights in its charts, however, lean on GLM-5.2 — a model that is not the current frontier.
| Benchmark | Beam | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|
| Terminal-Bench v2.1 | 80.1 | 88.3 | 90.6 |
| DeepSWE v1.1 | 44.4 | 68.0 | 74.2 |
| Humanity’s Last Exam | 36.2 | 46.9 | — |
| SWE-bench Verified | 80.9 | — | — |
| AIME 2026 | 97.8 | — | — |
Against the actual leaders in agentic coding, Beam loses every benchmark it was built to win. The Hacker News thread hit 500 points within hours, and the dominant reaction was pointed: Reflection built something expensive that underperforms cheaper alternatives on the benchmarks that matter most for its stated use case.
The efficiency argument is more defensible. If Beam matches GLM-5.2’s reasoning quality at a fraction of the inference cost, that is genuinely useful for production workloads where GLM-5.2 already cleared the quality bar. The caveat: Reflection’s “3-4x less compute” metric counts active parameters and output tokens, but excludes prompt processing, attention costs that grow with context length, and serving overhead. Real-world savings depend heavily on your workload mix.
The Catch: You Can’t Use It Yet
Beam is not available. Weights are promised under an Apache 2.0 license later in October 2026, alongside a technical report, model card, and tooling for running, evaluating, and fine-tuning. As of today, access is limited to a waitlist API. No pricing has been published. The API caps context at 262,144 tokens — not the 1 million claimed. All published benchmark scores come from Reflection itself; no third-party verification exists yet, and red-teaming is still underway.
When the weights do release, the hardware requirements are steep. FP8 serving needs roughly 500 GB of GPU memory — a full 8-GPU H200 or B200 node at minimum. This is not a model you run on a workstation. For most teams, the practical path runs through Reflection’s hosted API.
Who Built This and Why It Matters
Reflection AI’s founding team carries serious credentials. Misha Laskin led reward modeling for DeepMind’s Gemini project. Ioannis Antonoglou co-created AlphaGo. The company raised .63 billion total, including a .5 billion round in March 2026 at a 7.5 billion valuation. It signed a .3 billion compute deal with SpaceX for dedicated Nvidia GB300 NVL72 infrastructure. This is not a startup hoping to figure out the compute — it locked in resources before training began.
The geopolitical framing is deliberate. Chinese open models — Qwen 3.8 Max, GLM-5.2, DeepSeek, Kimi K3 — have dominated the open-weight frontier throughout 2025 and 2026. US government and enterprise procurement increasingly requires US-origin AI infrastructure, and Beam is positioned to capture that compliance demand regardless of raw benchmark performance.
What to Do Right Now
If you are building agentic pipelines and need a US-origin open-weight model with full Apache 2.0 commercial rights, Beam belongs on your evaluation list. Join the waitlist at api.reflection.ai and watch for the weight release this month. When it drops, benchmark against your specific workload before making infrastructure decisions — Reflection’s published numbers look better against last year’s competition than the current frontier.
If you need agentic coding capability today and don’t have US-origin compliance requirements, Kimi K3 and DeepSeek V4.1 Flash are the models Beam is measured against — and both are available now. Beam is a meaningful development for the US open-weight ecosystem. It is not, based on current data, a model that beats the state of the art.













