Reflection’s 501B Beam Is Announced—but the Weights Are Not Out Yet

|Author: QUASA Editorial Team|5 min read| 12
Reflection’s 501B Beam Is Announced—but the Weights Are Not Out Yet

SiliconANGLE’s October 5, 2026 launch report records Reflection AI’s announcement of Beam, a 501B-parameter language model initially offered through early access, with public weights planned for later in October. That timing leaves developers with a description of the model and a route to request access, but no generally downloadable checkpoint.

Beam is a sparse mixture-of-experts system built for coding, reasoning and agent tasks. It has 501 billion total parameters and 23 billion active for a given token; the larger figure describes the model as a whole, while the smaller one describes the part used during generation. Final red-teaming and evaluations remain underway. The planned public release includes the weights, a technical report, a model card and developer tools.

What the architecture and training figures show

Reflection’s technical announcement discloses pretraining on 23.8 trillion tokens from the web and licensed datasets, followed by more than 100 million reinforcement-learning rollouts on about 10,500 NVIDIA GB300 GPUs over four weeks. These are company-reported training figures, rather than the contents of a reproducible training package. They establish the scale of the work behind Beam while leaving the precise evaluation and deployment details to the promised report and model card.

Sparsity is central to the efficiency argument. In a mixture-of-experts model, routing selects a subset of parameters for each token, so the full parameter count and the amount of computation used to generate a token answer different questions. The active count can help explain why a model with a very large stored weight set may require less generation compute than its total size implies. It does not tell a team what a real deployment will cost: memory, prompt processing, attention and serving choices remain relevant.

The training account separates a base-model phase from later work aimed at reasoning and tool use. Pretraining supplied code, technical material and other web or licensed content; midtraining extended the stated effective context length to one million tokens. That is a training claim, not a published serving limit for every eventual Beam deployment. Reinforcement learning then used coding, agent and STEM environments, with rollouts generated and graded in sandboxes.

What developers can access now

Announcement and early access: A public description, benchmark tables and an early-access waitlist are available. An early version has been made available to a select group of users, but joining the waitlist is a request rather than immediate access. The difference matters for teams trying to schedule an evaluation: a published score or demo is an account of a test, while access to an endpoint or weights permits work on their own tasks.

Weights and license: The stated plan is to publish the weights under Apache 2.0 later in October. Until a checkpoint and its license text are available, the license describes the intended terms of a future release. The public waitlist does not give applicants the ability to self-host, modify or redistribute Beam’s weights. The actual file set, formats and distribution arrangements will matter when that plan becomes a release.

Documentation and safety: The technical report, model card and artifacts for running, evaluating and fine-tuning Beam are scheduled to accompany the weights. The safety work involved a separate model trained from the pretrained checkpoint and combined with the main model through distillation. Safety evaluation results are due in the technical report, and internally developed safety evaluations are also slated for open release. Those materials would give prospective users a more detailed account of methods, limitations and behavior than the launch tables alone.

What the benchmark comparison measures

Help Net Security’s benchmark breakdown cites the vendor’s Terminal Bench v2.1 scores: 80.1 for Beam, 81.0 for GLM 5.2, 88.3 for Kimi K3 and 90.6 for DeepSeek V4.1 Flash. On that command-line task comparison, Beam sits just below GLM 5.2 and farther behind the other two named models. These numbers come from a launch table, not an independent rerun of Beam’s public weights.

The table contains different tasks and different gaps. On DeepSWE v1.1, the listed Beam score is 44.4 against GLM 5.2’s 44.0 and DeepSeek V4.1 Flash’s 74.2. A score on one coding test cannot stand in for all coding, reasoning or agent use, particularly when some comparison cells are marked as unreported. Independent evaluations will become possible once the checkpoint and accompanying documentation are available.

The accompanying efficiency comparison estimates generation forward-pass computation by multiplying the active parameter count by the average number of generated tokens and an operation-count factor. It omits prompt prefill, context-dependent attention and serving overhead. That makes the comparison a way to frame generation work under the chosen method, rather than a measured inference bill or a guarantee of lower cost for any particular deployment. Longer reasoning traces could also change the token side of the calculation even if the active parameter count stays fixed.

The next substantive event is the promised October release of the weights, technical report and model card. Once those artifacts appear, developers can inspect the shipped configuration, examine the license attached to it and run evaluations that match their own workloads. The timing of that release will determine when Beam’s open-weight plan becomes independently testable.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0