Reflection is accepting early-access sign-ups for Beam, a 501-billion-parameter LLM built for coding, reasoning, and agentic workloads. The company describes it as its first open-weight foundation model, though downloadable weights are still pending.
In its Beam announcement, Reflection says the model is undergoing final red-teaming and evaluations. It promises the weights, technical report, model card, and developer artifacts later in October 2026. For now, developers have an announcement and an early-access invitation.
Beam was announced on October 5, offering another prospective Western option in a market with strong Chinese open-weight competitors. Its pitch combines competitive capability with lower inference compute. The evidence so far consists primarily of Reflection’s published results and methodology. As TechCrunch’s launch coverage notes, the performance claims have not been independently verified.
Early Access Is Not the Open-Weight Release
At announcement time, Beam’s weights, model card, technical report, license terms, and developer artifacts remained unpublished. Interested users can sign up for early access, but that does not guarantee immediate access, public API availability, or a downloadable model. Reflection has established what it intends to publish; the conditions under which developers can use it remain unknown.
The promised October artifacts serve different purposes. Weights would make it possible to run the model outside Reflection’s own environment, subject to the eventual license and software support. A technical report should provide a fuller account of the architecture, training, and evaluations, while a model card should help readers assess intended uses and limitations.
Enterprises need the license terms to determine which commercial uses or redistribution arrangements are permitted. Infrastructure teams need deployment documentation to reliably plan hardware, serving software, and integration work.
“Open-weight” is not a synonym for unrestricted use or a fully open development process. The eventual license and accompanying disclosures will determine what Reflection’s commitment means in practice.
Beam Activates 23B of Its 501B Parameters
According to Reflection, Beam uses a sparse mixture-of-experts architecture with 501 billion total parameters and 23 billion active parameters.
A mixture-of-experts model routes tokens through selected expert components instead of activating the entire network for every token. The total parameter count describes the model’s overall size, while the active count is more relevant to part of the computation performed during generation. This distinction helps explain Reflection’s efficiency argument.
Beam is not equivalent to a conventional 23-billion-parameter model, however. Storing and serving the larger collection of weights still creates infrastructure requirements, and Reflection has not yet supplied enough deployment information to turn the active-parameter figure into a reliable hardware recommendation.
Reflection describes Beam as text-only. Its demonstrations include tasks involving visual information and documents, but the company explains that the model can use tools and work with other modalities when their information is represented as text.
An agent calling an OCR service to obtain document text is different from a model natively interpreting an image. Beam’s tool-assisted demonstrations therefore should not be mistaken for evidence of built-in multimodal capability.
The Benchmark Table Shows Strengths and Gaps
Reflection’s published benchmark table covers software engineering, terminal tasks, reasoning, search, tool use, and general capabilities. The results offer a more qualified picture than a claim that Beam leads open-weight AI models across the board.
On SWE-bench Verified, the company reports 80.9 for Beam, compared with 77.6 for Inkling and 70.7 for Nemotron 3 Ultra. Several other models have no reported result in that row, so it cannot support a ranking against every competitor listed elsewhere.
Beam’s reported 80.1 on Terminal Bench v2.1 is close to GLM 5.2’s 81.0 but below Qwen 3.8 Max’s 86.6 and Kimi K3’s 88.3. Reasoning results show gaps as well: Beam scores 36.2 on Humanity’s Last Exam without tools, against 40.5 for GLM 5.2 and 46.9 for Kimi K3.
These comparisons are vendor-published; Beam’s results have not been independently reproduced. Reflection says it uses Artificial Analysis and DataCurve as sources for other models’ evaluations, which does not mean those organizations independently validated Beam.
The table supports Reflection’s positioning of Beam as a potentially useful coding and agentic model with an efficiency emphasis, without establishing universal superiority. For agentic benchmarks, tool access, context management, reasoning settings, and the number of attempts can all affect results. The technical report will be important for determining how directly comparable the published scores are.
Lower Estimated Compute Does Not Establish Lower Cost
Reflection’s most consequential claim may be that Beam reaches scores comparable to GLM 5.2 on advanced reasoning benchmarks while using three to four times less inference compute.
The company’s compute-estimation methodology estimates generation forward-pass computation using:
FLOPs ≈ 2 × active parameters × mean generated tokens per attempt
FLOPs are floating-point operations, a measure of computational work. The estimate counts both reasoning tokens and final-answer tokens. For mixture-of-experts models, it uses the parameters activated per token instead of the full model size.
This is a useful way to examine how model size and response length interact. A model that activates fewer parameters and reaches an answer with fewer tokens can require less estimated generation compute, even if it does not produce the highest benchmark score.
Reflection explicitly excludes prompt prefill, context-dependent attention operations, and serving overhead. Its calculation is an approximate compute comparison, not a measurement of inference cost.
Readers therefore cannot infer price, latency, or infrastructure efficiency directly from the estimate. Actual deployment economics also depend on hardware utilization, memory requirements, batching, and the serving implementation. A lower estimate does not establish that a particular provider will offer a proportionally cheaper API or that a self-hosted deployment will deliver the same savings.
Users can adjust Beam’s reasoning effort, according to Reflection, with lower settings favoring shorter responses and higher settings allowing longer reasoning. That could give developers a useful capability-cost control. Its practical value depends on how much task success changes at each setting, alongside the number of tokens saved.
Reflection Attributes the Gains to Large-Scale RL
Reflection says it pretrained Beam on 23.8 trillion tokens drawn from the web and proprietary licensed datasets, then made reinforcement learning, or RL, a central part of capability development.
Its high-compute RL campaign reportedly used 10,500 NVIDIA GB300 GPUs over four weeks, generating more than 100 million rollouts. A rollout is a model-generated attempt or interaction sequence used in the training process.
The company also reports approximately 1.3 billion sandboxes for training and grading, and one million coding, agentic, and STEM environments. These substantial vendor-reported figures have not been independently audited. They describe the claimed scale of the campaign without proving the resulting model’s quality.
Sources
- Beam announcementreflection.ai
- TechCrunch’s launch coveragetechcrunch.com





