Reflection's Beam: 501B Parameters, 23B Active. Cheaper to

Abstract tech illustration: Reflection's Beam: 501B Parameters, 23B Active. Cheaper to

Reflection AI says its new open-weight model, Beam, delivers GLM-5.2-class coding and reasoning at three to four times less compute. That sounds like a smaller API bill, but compute estimates, per-token prices, and cost per finished job are three different numbers, and so far only the first exists. Here's what Reflection actually claimed, what nobody has verified yet, and how to test it on your own workflow when the weights land.

What Reflection actually shipped (and what it didn't)

Beam is Reflection's first open-weight model, announced October 5, 2026. It's a sparse Mixture-of-Experts with 501B total parameters and 23B active per token, aimed at coding, reasoning, and agentic workloads (Reflection's announcement).

Three things to separate:

  • Weights are not out yet. Reflection says it will release them under Apache 2.0 later in October 2026, along with a technical report, model card, docs, and the full stack for running, evaluating, and fine-tuning. Per MarkTechPost, the model is in final red-teaming and early access is a waitlist on Reflection's platform. As of the sources I read, you can't self-host it today. Check Hugging Face before you plan around it.
  • There's no published per-token price. DataCamp's writeup notes that as of October 6, a dollar comparison against hosted Kimi K3 or DeepSeek V4.1 Flash endpoints wasn't possible. Any "Beam is cheaper" claim in dollars is unproven.
  • Every benchmark score is self-reported. Trending Topics points out that no independent benchmarks exist, because the weights aren't public yet and nobody outside Reflection can run the model.

Treat everything below as a claim to be tested.

What the 3–4x efficiency number measures

The 3–4x figure is an estimate of generation compute, not a measured inference cost. According to Kingy AI's breakdown, it's calculated as:

compute ≈ 2 × active_parameters × mean_generated_tokens_per_attempt

Generated tokens include both reasoning and the final answer. FourWeekMBA adds that the estimate leaves out prompt prefill, context-dependent attention operations, and serving overhead, and that Reflection itself calls it an approximate comparison.

Why that matters for a small business: your real workloads are often prompt-heavy. Think of an invoice extraction job where you feed in a long PDF and get back 200 tokens of JSON, or a support ticket classifier reading a long thread. For those jobs, the part the formula leaves out (prefill) may dominate your cost. The 3–4x number was also stated for advanced reasoning benchmarks, which look nothing like "categorize this ticket."

The comparison model also matters. GLM-5.2 is a MoE with roughly 750B total and about 40B active parameters (Codersera; sources disagree slightly on exact counts). Beam's 23B active is smaller, but the 3–4x claim also depends on how many tokens each model generates per attempt. A model that thinks longer burns the savings back.

What the benchmark table shows

By Reflection's own reported scores (Implicator), Beam and GLM-5.2 are close:

Benchmark Beam GLM-5.2
Terminal Bench v2.1 80.1 81.0
DeepSWE v1.1 44.4 44.0
SWE Bench Pro v1 65.5 62.1

Beam is slightly behind on one, slightly ahead on two. "GLM-5.2-level" is a fair summary of these three numbers. It's not a claim of across-the-board superiority.

It also isn't the strongest open model. In Reflection's own table, Beri.net reports Beam trails Kimi K3 on 13 of the 14 benchmarks where both have a score. On Terminal Bench v2.1, Kimi K3 is listed at 88.3, GLM-5.3 at 88.2, and DeepSeek V4.1 Flash at 90.6, against Beam's 80.1. So the pitch is efficiency and provenance, not raw capability. If you just want the highest score on an open model, Beam isn't it.

The memory footprint is the catch

Active parameters determine compute per token. Total parameters determine how much hardware you need to hold the model. Beam is 501B total, so a self-hosted deployment has to keep all of it in memory.

One analysis (AlphaSignal) estimates raw weights at about 501 GB in FP8 or about 251 GB in 4-bit, before KV cache and overhead. Reflection plans FP8 and NVFP4 downloads.

Rough translation, as illustrative math: a model that needs 250+ GB for weights alone doesn't fit on a workstation GPU or a Mac mini. You're renting a multi-GPU node, or you're calling someone else's hosted endpoint. For a 1-10 person company, that changes the question from "can I self-host it?" to "will a provider host it at a price that beats what I pay now?" And on that question, Beam has no published price yet.

For comparison, GLM-5.2 has open weights under the MIT license (VentureBeat), and Z.ai's official API lists $1.40 per 1M input tokens, $0.26 per 1M cached input, and $4.40 per 1M output (Z.ai pricing docs). Prices vary by provider and change often, so re-check before you budget. That's the number Beam has to beat once someone hosts it.

Why a US-built open model matters even if it isn't the best

The real argument for Beam isn't the benchmark table. It's who's allowed to use it. For clients with procurement rules, regulated data, or a cautious legal team, an open-weight model from a US lab removes an objection that has nothing to do with quality. It doesn't make Beam better. It widens the set of businesses that can say yes.

The second benefit is leverage. If your automations are wired to one vendor's API, every price hike and rate limit lands on you. With open weights under a permissive license (Apache 2.0, if the published license file matches what Reflection announced; verify that when the weights drop), you can run the same model at more than one provider and walk if the price moves. Open weights don't mean cheap, though. Hosting a model this size has its own bill, and a hosted API from a bigger lab can still win for low-volume work.

How to test Beam on your own work in an hour

Once Beam is on a provider you can call, skip the leaderboards and run this:

  1. Pull 20 real tasks from your business: client emails you've drafted replies to, invoices you've extracted data from, support tickets you've categorized.
  2. Run the same 20 through your current model and through Beam.
  3. Score one thing: did it finish the job without me fixing it? Pass or fail.
  4. Compute cost per finished job, not cost per token.

A minimal harness looks like this. The point is that the model is a config value, not something hardcoded in fifty places:

import os, json, time
from openai import OpenAI  # any OpenAI-compatible endpoint

MODELS = {
    "current": {"base_url": os.environ["CURRENT_URL"], "model": os.environ["CURRENT_MODEL"],
                "in_per_m": float(os.environ["CURRENT_IN"]), "out_per_m": float(os.environ["CURRENT_OUT"])},
    "beam":    {"base_url": os.environ["BEAM_URL"],    "model": os.environ["BEAM_MODEL"],
                "in_per_m": float(os.environ["BEAM_IN"]),    "out_per_m": float(os.environ["BEAM_OUT"])},
}

def run(name, tasks):
    cfg = MODELS[name]
    client = OpenAI(base_url=cfg["base_url"], api_key=os.environ[f"{name.upper()}_KEY"])
    results = []
    for t in tasks:
        r = client.chat.completions.create(
            model=cfg["model"],
            messages=[{"role": "user", "content": t["prompt"]}],
        )
        u = r.usage
        cost = (u.prompt_tokens * cfg["in_per_m"] + u.completion_tokens * cfg["out_per_m"]) / 1_000_000
        results.append({"id": t["id"], "output": r.choices[0].message.content, "cost": cost})
    return results

# After you hand-score each output as pass/fail:
# cost_per_finished_job = total_cost / number_of_passes

Fill the price variables from your provider's actual rate card on the day you test. Don't use numbers from a blog post, including this one.

The step people skip is the last one. If a cheaper model needs a human to fix a third of its output, your cost per finished job goes up, not down. Illustrative round numbers: if Model A costs $0.02 per call and passes 95% of the time, and Model B costs $0.01 per call but passes 60%, then A costs about $0.021 per finished job and B costs about $0.017 before you count your own time. Add two minutes of human cleanup on each failed job and B loses badly. Your time is the most expensive token in the system.

Also log latency and the length of the output, since a model that reasons longer can eat the compute savings in the formula above.

What I'd do this week

  • Don't switch anything yet. No independent evals, no price, no downloadable weights. Wait for the technical report and for at least one provider to publish a rate card.
  • Make your model a setting. If swapping models means editing prompts in a dozen workflows, you can't run the test above without a weekend of pain. One config file, one client wrapper, one place to change the model name.
  • Keep a standing 20-task test set. It's reusable for every model release, so you stop reading launch posts and start reading your own numbers.
  • Watch for the unglamorous details: the actual license file (confirm it's plain Apache 2.0), which providers host it, whether it shows up on OpenRouter, and what the FP8 serving requirements are in practice.

My view: efficiency is the story that matters in open-weight models, and the leaderboard race is mostly noise for small teams. Nobody running a ten-person agency needs the single smartest model available. They need one that's good enough, predictably priced, and not going to vanish or triple in cost. Beam might be that model. Its own numbers say it's competitive with GLM-5.2 and behind the newest open models, and the cost advantage is, for now, an estimate of generation compute. The only way to turn that estimate into a decision is to measure cost per finished job on your own workload.


Want more like this?

I write about practical AI automation and GenAI engineering for solopreneurs and small teams.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn — Lazar Milićević.

Visit bizflowai.io for services and case studies.

Frequently asked questions

What is Beam, Reflection's open-weight model?

Beam is Reflection's first open-weight model. It uses a mixture-of-experts design with 501 billion total parameters, but only about 23 billion are active for any given token. Reflection says it aims to match GLM 5.2 on coding and reasoning while using three to four times less compute. It is billed as the most capable open-weight model built outside China. These performance figures come from announcement coverage and are unverified claims.

Why does a mixture-of-experts model like Beam matter for small business costs?

A mixture-of-experts model only does the math for a slice of its parameters on each token. Beam activates about 23 billion of its 501 billion parameters, so per-token compute drops, which can lower inference costs whether you use a hosting provider or your own hardware. The catch: you still need enough memory to hold all 501 billion parameters, so Beam is not a laptop model. It suits rented GPU servers or hosted endpoints.

Why does it matter that Beam is an open-weight model from a US lab?

Most strong open-weight models currently come from Chinese labs, which can be a blocker for organizations with procurement rules, regulated data, or cautious legal teams. A capable open-weight model from a US lab removes that objection and widens who is allowed to use it. It does not make the model better, only more broadly usable. Open weights also let you host the model in several places and switch if prices change.

How do I test whether Beam is worth switching to for my business?

Skip benchmarks and test your own work in about an hour. Pull twenty real tasks, such as client email replies, invoice data extraction, or support ticket categorization. Run them through your current model and through Beam on any provider that hosts it. Score each on whether it finished the job without your fixes, then compare cost per finished job rather than cost per token. A cheaper model that needs human fixes on a third of its output isn't cheaper.

When should I use an open-weight model like Beam vs. a hosted API from a bigger lab?

Open weights give you leverage: you can host the model in more than one place and walk away if a price moves. But open weights don't automatically mean cheap, because hosting a model this size has its own bill. A hosted API from a bigger lab can still win on price at low volume. Test both on your real workload, and build automations so the model is a swappable setting rather than hardcoded.