Reflection AI's Beam, published Oct. 5, is a 501-billion-parameter open-weight model the Brooklyn startup says reasons at a level comparable to Z.ai's GLM-5.2 while using 3 to 4 times less inference compute. The model activates only 23 billion of its parameters per token, a sparse mixture-of-experts design pretrained on 23.8 trillion tokens before its context window was extended to 1 million tokens in a later training stage. None of it is independently confirmed yet: Reflection ran every benchmark itself, and the efficiency figure is the company's own estimate, not a measurement.
On the numbers Reflection chose to publish, Beam scores 80.9 on SWE-Bench Verified, 97.8 on AIME 2026 and 90.5 on GPQA Diamond -- strong results that, if they hold up outside the company's own test harness, would make Beam one of the most capable fully open models available under a permissive license. Reflection says weights, a technical report and a model card will ship under an Apache 2.0 license before the end of October; for now, only a small early-access group can run Beam at all.
The training run behind those numbers was itself a large infrastructure undertaking: Reflection says pretraining ran on 6,144 Nvidia GB300 NVL72 GPUs in under four weeks, reaching 92.3% "goodput" toward completing the run, while the reinforcement-learning stage used 10,500 GB300 GPUs over four more weeks to generate over 100 million rollouts across nearly one million distinct training environments. The company also says it handled 71 infrastructure incidents during training with only 0.02% of total GPU-minutes lost to them -- an operational-reliability claim, like the benchmark scores, that comes from Reflection's own account of its own run rather than an outside audit.
What shipped Oct. 5 -- and what didn't
- Parameters
- 501B total, 23B active per token (sparse MoE)
- Training data
- 23.8 trillion tokens; context extended to 1M tokens later
- Headline claim
- Comparable to GLM-5.2 at 3-4x less inference compute (Reflection's own estimate)
- License
- Apache 2.0, weights due before end of October 2026
- Independent verification
- None yet -- Artificial Analysis has early access, no published score
The efficiency claim is the headline, and it rests on math Reflection built itself. The company estimates inference compute as roughly twice Beam's active-parameter count multiplied by the average number of tokens a response generates -- a calculation that excludes prompt processing, attention operations and the overhead of actually serving a model to real traffic. Reflection calls this an approximate comparison rather than measured inference cost, which is itself more candid than most vendor efficiency claims in this market, even as the underlying number stays unverifiable from outside.
What the "3-4x less compute" claim covers -- and doesn't
- 3-4x · less inference compute vs. GLM-5.2
- Reflection's own estimate, not a measurement
Includes: Roughly 2x active-parameter count x mean generated tokens, by Reflection's stated method
Excludes: Prompt/prefill processing, attention operations, real-world serving overhead, and any independent benchmark run
GLM-5.2 is Z.ai's flagship open-weights model: roughly 744 billion total parameters with about 40 billion active per token, released in June 2026 and widely adopted since. It's a reasonable yardstick -- both models are openly licensed, and both target coding and agentic workloads. But picking a June rival is also a choice. Z.ai has since shipped GLM-5.3, and Beam's own benchmark page shows it trailing that newer model, trailing Moonshot AI's Kimi K3, and trailing DeepSeek V4.1 Flash on most coding tasks -- none of which made it into the launch's headline comparison.
Beam vs. the model it headlined against
| Beam | GLM-5.2 | |
|---|---|---|
| Total / active parameters | 501B / 23B | ~744B / ~40B |
| License | Apache 2.0 (pending) | Open-weights |
| Release date | Weights due late Oct. 2026 | June 2026 |
| Independently verified | No -- self-reported | Yes, measured since release |
That's the reconciliation Reflection's own launch doesn't volunteer: Beam is being marketed against a model it can match, not the models that have since passed it. The gap matters less for what it says about Beam's raw capability -- a model that's merely competitive with a six-month-old rival is still a serious release -- than for what it says about how to read any benchmark claim at launch. The comparison a company chooses to headline is itself a decision, and this one picked the fight it could win.
The launch lands inside a fast-moving financial story of its own. Reflection AI, founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, has raised roughly $4.7 billion to date, including a round that valued the company at $25 billion pre-money -- a jump from the $545 million valuation it carried seven months earlier. Nvidia, Sequoia Capital and Lightspeed Venture Partners are among the backers, and Reflection has separately signed multibillion-dollar compute-capacity deals with Nebius and SpaceX running through 2029.
Reflection AI pitches Beam less as a chatbot competitor than as a foundation for what it calls "AI factories" -- enterprises and governments training customized systems on their own proprietary data, distributed through cloud providers and open-source integrations rather than a single API Reflection controls. The pitch targets buyers wary of building on Chinese open models but unwilling to depend on a closed Western lab's terms either -- a narrower, more specific market than "everyone who uses AI."
- Beam reasons at a level comparable to GLM-5.2 while using 3-4x less inference compute.
- Beam is one of the strongest fully open models available on coding and reasoning benchmarks.
- Beam trails GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash on most coding tasks.
None of this makes Beam a bad model -- it makes it an unverified one, which in this market is the default state of every launch-day benchmark claim, not a special flaw. The real test starts when Artificial Analysis publishes an independent score, and when the Apache 2.0 weights are actually in outside hands at the end of the month. Until then, the honest read of Beam is Reflection's own: fast, open and, by the company's own fuller benchmark page, not yet the fastest open model available -- even to the company that built it.
Beam is being marketed against the model it can match, not the models that have since passed it.
- Reflection AI published self-run benchmark scores for Beam, a 501-billion-parameter open-weight model, on Oct. 5.
- The company claims GLM-5.2-level reasoning at 3-4x less inference compute -- its own estimate, not an independent measurement.
- Backed by Nvidia at a $25 billion valuation, Reflection pitches Beam as sovereign AI infrastructure for enterprises and governments.
- By its own benchmark page, Beam trails newer rivals GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash on most coding tasks.
- Caveat: Artificial Analysis has early access but hasn't published results, so every number in this piece is still self-reported.