Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Reflection AI unveils Beam, a 501B-parameter (23B active)…

Reflection AI unveils Beam, a 501B-parameter (23B active) Apache-2.0 open-weight MoE; weights due later in October

★★★★after cutoffmodel-releaseReflection AIconfidence: high

On Oct 5, 2026 Reflection AI announced Beam, its first model: a sparse Mixture-of-Experts with 501B total and 23B active parameters, trained on 23.8T tokens, for coding, reasoning and agentic work. It will be released under Apache 2.0 with weights and a tech report "later in October"; for now it is in early access through Reflection's beta API. Reflection says it matches Z.ai's GLM-5.2 on reasoning with 3-4x less inference compute and beats Thinking Machines' Inkling and Nvidia's Nemotron 3 Ultra, but its own table shows it trailing GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash on most tests.

Key facts

What happened

A day after Axios reported that a release was imminent, Reflection AI published "Introducing Beam" on Oct 5, 2026. Beam is the lab's first model and is trained from scratch (not distilled, per the founders' interview with Sources). It is a sparse MoE with 501B total and 23B active parameters, aimed at coding, reasoning and agentic workloads. It was trained on 23.8T tokens, with pretraining on 6,144 GB300 GPUs in under four weeks and a four-week RL run on 10.5K GB300 GPUs that Reflection calls "one of the largest scale RL runs conducted by any open lab to date."

The weights are not out yet. Reflection says Beam is in "final red-teaming and evaluations" and that the weights and a technical report will follow later in October under Apache 2.0. Until then a waitlisted early version runs on Reflection's beta API (Beam-501B-A23B, OpenAI-compatible endpoint). Distribution through hyperscalers, neoclouds and open-source libraries is planned (TechCrunch).

Reflection's pitch is efficiency: GLM-5.2-level reasoning at 3-4x less inference compute (about a quarter of GLM-5.3's, per Sources), and "approaching" Qwen 3.8-Max on coding and agentic tasks (Semafor). The blog's own table is mixed. Beam leads the Western open models it lists (Thinking Machines' Inkling, Nvidia's Nemotron 3 Ultra) on most coding benchmarks, but it trails GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash on Terminal Bench, HLE, SciCode and AutomationBench. Hacker News commenters said the charts leave out those stronger rivals. CEO Misha Laskin compared closed models to "renting an apartment" (Sources).

Why it matters

This is the first model from the best-funded US lab dedicated to open weights (about $4.6B raised, $25B valuation, more than $7B in GPU deals). An Apache-2.0 500B-class US model gives enterprises and governments that avoid Chinese weights a near-frontier option. The benchmarks still put it behind the best Chinese open models, so it narrows the US–China open-weight gap without closing it. Watch for the actual weight release, the tech report and independent evaluations.

Changelog

  • 2026-10-05: created (official blog and developer docs; TechCrunch, Semafor, Sources; Bloomberg paywalled, headline only)

Models

People

Ioannis Antonoglou Misha Laskin

Related events

  1. Axios: Nvidia-backed Reflection AI is about to release its first open-weight model, with other Western open-weight models due in October ★★★
  2. Thinking Machines Lab releases Inkling, its first open-weights model (975B MoE) ★★★★
  3. Aleph Alpha releases Kolibri, a 78B-parameter (3.5B active) open-weight English–German 'sovereign' MoE model under Apache 2.0 ★★★
  4. Open-weight models carry a majority of tokens on Vercel's AI Gateway for the first time (56% in August 2026) ★★★

Sources (8)

id: 2026-10-05-reflection-beam-501b-open-weight · updated 2026-10-05 · open in the interactive timeline