Industry NewsIndustry News 6 min read

Reflection Unveils Beam, a 501B Open-Weight Model

Reflection AI unveiled Beam on 5 October 2026, a 501-billion-parameter open-weight model that it says matches GLM-5.2 on reasoning with three to four times less inference compute.

PC

PromptCrates Editorial

Staff Writer

0 0
Reflection Unveils Beam, a 501B Open-Weight Model

Reflection AI on Monday, 5 October 2026, unveiled Beam, its first open-weight model, a sparse mixture-of-experts system with 501 billion total parameters, of which 23 billion are active for each token. The Brooklyn-based startup says Beam matches Z.ai’s GLM-5.2 on advanced reasoning benchmarks while using three to four times less inference compute, and it plans to release the weights under an Apache 2.0 license later this month.

The launch, laid out in a long Reflection blog post, is the most direct attempt so far by a US startup to compete with the Chinese labs that dominate the top of the open-weight rankings. Beam is still in final red-teaming and evaluation, and for now it is available only to a select group through an early-access waitlist.

What Reflection says Beam can do

Reflection pitches Beam as a workhorse for coding and agent tasks rather than the most capable model on every chart. It was pretrained on 23.8 trillion tokens from the web, public sources and licensed datasets, supports a context window of one million tokens, and is text-only, although Reflection says it can handle information from other modalities once that information is expressed as text. A reasoning effort setting lets users trade shorter answers for deeper reasoning on harder problems.

The company’s own benchmark tables show a model that is competitive rather than dominant. Beam scores 80.9 on SWE-Bench Verified, ahead of Thinking Machines Lab’s Inkling at 77.6 and Nvidia’s Nemotron 3 Ultra at 70.7. On Terminal Bench v2.1 it posts 80.1, just under GLM 5.2 at 81.0. It reaches 90.5 on GPQA Diamond against 91.2 for GLM 5.2, and 36.2 on Humanity’s Last Exam without tools, where GLM 5.2 scores 40.5 and Kimi K3 46.9. Reflection openly concedes that frontier open models such as Kimi K3 remain ahead on raw capability and argues that Beam’s edge is efficiency at inference time.

TechCrunch stressed that none of these performance claims have been independently verified, and pointed out that Inkling, released in July, is multimodal while Beam is not. Reflection’s demos include a live New York subway map built from public transit data and a fine-tuning notebook for a small Gemma-4 model that, by the company’s account, lifted that model’s held-out accuracy on a Text2SQL task by 66.5%.

The compute behind Reflection’s first frontier model

The engineering details are the most striking part of the post. Reflection says it pretrained Beam end to end in under four weeks on a cluster of 6,144 Nvidia GB300 GPUs, with training goodput reaching 92.3% toward the end of the run and nine semi-automatic rewinds along the way.

Reinforcement learning was the bigger bet. The company says it ran 10,500 GB300 GPUs for four weeks, generating more than 100 million rollouts with contexts of up to 256,000 tokens, and used about 1.3 billion sandboxes for training and grading. It sourced nearly one million coding, agentic and STEM environments and says capabilities kept improving with more RL compute, with no sign of a plateau. During the run its systems handled 71 inference incidents without stopping training, recovering capacity in a median of eight minutes.

That hardware did not come cheap. SiliconANGLE said Reflection reportedly signed a $6.3 billion deal with SpaceX to rent the GB300 NVL72 systems it used for training, while TechCrunch said compute deals with SpaceX and Nebius signed this summer are collectively worth more than $7 billion and secure access to GB300 chips through 2029.

On safety, Reflection says it trained a separate safety and alignment model from the same pretrained checkpoint and merged it with the main reinforcement learning model through multi-teacher on-policy distillation. It promises to publish safety evaluation results in the technical report and to open-source the safety tests it built internally.

Why Beam matters in the open-weight race

Reflection has the money to keep going. TechCrunch reported that the company, founded in 2024 by two former Google DeepMind researchers, has raised roughly $4.7 billion from backers including Nvidia, Sequoia Capital and Lightspeed Venture Partners, according to PitchBook, and that its last round valued it at $25 billion before the new money.

Its business plan targets enterprises and sovereign governments. TechCrunch described the pitch as AI factories, where institutions train Reflection’s models on their own proprietary data to build customized local systems, and said the startup has begun testing a sovereign AI factory partnership with Shinsegae Group in South Korea. SiliconANGLE called Beam the first open model from a US startup to show comparable or better performance than leading Chinese models such as GLM-5.2 and Qwen 3.8-Max, while noting that free models still trail closed frontier systems like Anthropic’s Claude Fable 5.1.

The context helps explain the excitement. PromptCrates has covered how Chinese labs keep pushing open weights, including when Z.ai confirmed it built Ox Alpha with open weights to follow, and how open-weight AI firms became the hottest acquisition targets this summer. A credible US-built alternative gives companies that are wary of Chinese-origin models a new option to evaluate.

What developers should watch before the weights land

For teams deciding whether to plan around Beam, the timing matters more than the benchmark tables. Until the weights, model card and technical report arrive, outside researchers cannot check the efficiency claims or run the model on their own hardware. Reflection says the release will come with documentation, the full stack for running, evaluating and fine-tuning the model, distribution through partners and integrations with popular open-source libraries and harnesses.

Running a 501-billion-parameter model is still a data center job, even with only 23 billion parameters active per token, so most developers will reach Beam through hosted endpoints rather than local machines. Smaller open models, such as the Qwen3.8-27B release that put local agents on a desktop, remain the practical choice for laptops. Reflection says it is already training the next model in the series. The real test will be whether independent evaluations confirm Beam’s efficiency once the Apache 2.0 weights are public.

industry-newsReflection AIopen-weight modelsLLMNvidia GB300

Related articles