
What is Reflection AI Beam?
Beam is the first model from Reflection AI, a lab that two former Google DeepMind researchers started in 2024. Under the hood it's a sparse mixture-of-experts (MoE) model, with 501 billion parameters in total, 23 billion active on any given token. The training target was coding and reasoning work, with agent workloads on top of that.
That puts it in the same lane as Kimi K3, Qwen3.8 Max, and DeepSeek V4 Flash as well.
It also sits next to the other US open-weight bet of this year, Inkling.

I build AI agents for a living, so I tend to read a launch post like the spec sheet of a car before the test drive, where the numbers that matter are the ones that end up on your bill and in your latency logs. With Beam, what makes the numbers interesting is mostly how they were made:
- Pretraining:
23.8Ttokens, run end to end in under four weeks on6,144NVIDIA GB300 GPUs, per Reflection's launch post. - Reinforcement learning:
10.5KGB300 GPUs for four weeks, generating more than100 millionrollouts across nearly one million training environments. - Architecture:
52layers, interleaved local and global attention, and fine-grained routed experts with near-uniform utilization (the busiest expert runs at just1.04xaverage load). - Input: text only. Image, audio and file inputs are not supported on the API, per the OpenAI compatibility docs.
- License: Apache 2.0, promised "this month" alongside quantized FP8 and NVFP4 versions, per Reflection's launch thread.
The launch got a lot of attention. The official post on X passed 8,492 likes, while the Hacker News thread reached 549 points and 174 comments in its first days.
What's usable today vs what was announced
Most launch coverage skips this part, which is a shame, because it's the part that decides if Beam belongs in your plans this month at all. The announcement is describing the model Reflection trained, and the developer docs are describing the model you can actually call. For now, those two are not the same thing.

| What the launch says | What the API and docs say (Oct 8, 2026) | |
|---|---|---|
| Weights | Apache 2.0 "this month" | Reflection's Hugging Face org has 0 models |
| Context window | Effective context extended to 1M tokens | 256K context, 128K max output, "may change during beta" (Models) |
| Access | Early access for "a select group of users" | Waitlist; keys only once access is enabled (Authentication) |
| Price | "Lower cost" via efficiency | No public rate; the pricing page sits behind login |
| Rate limits | Not mentioned | Per-organization limits and a daily token allowance, no numbers published (Rate limits) |
| Reasoning | Controllable effort | Always on, five levels from low to max, default medium (Reasoning) |
| Knowledge cutoff | Not stated | June 30, 2026 |
The context window is the gap I'd flag hardest. If your plan is a long-document or whole-repo workflow built around "1M tokens," the beta API today caps you at a quarter of that. I wouldn't call it a broken promise, since training capability and serving limits are two different things, but it does change what you can build in October.
Hacker News user adrian_b spelled out the deadline in plain terms:
"They claim that they will release the model as open weights later this month. That means that they have the 31th of October as the deadline to make true their claims."
How Beam works: 23B working, 501B stored
The MoE design is pretty much the whole story of Beam, so it's worth spending one plain-language minute on it. A mixture-of-experts model splits its weights into many small "expert" networks, and for each token a router picks a handful of them, then only those picked experts do any math. This is the reason a 501B model can cost about as much per token to run as a 23B dense model does.

The catch that people tend to miss is this one: active parameters set your compute bill, total parameters set your hardware bill. All 501B still need to sit in memory, because on the next token the router can pick any expert. So per token Beam is cheap to run, but it's not small to host. Next to Kimi K3 at 2.8T total and 104B active, Beam is a lighter lift on both counts, and next to DeepSeek V4 Flash, which activates only 13B, it does more work per token.
Reasoning length is the other lever here. Reflection trained Beam with a length penalty that rewards correct answers while discouraging tokens it doesn't need, then exposed that tradeoff through the reasoning_effort parameter. In their efficiency charts, the score is plotted against how many tokens each model spends:

Under the chart, Reflection states one honest caveat itself: the compute estimates leave out prompt prefill and attention costs, as well as serving overhead, so they "represent an approximate compute comparison rather than measured inference cost." Put another way, what you see is the shape of the efficiency and not a bill.
Beam benchmarks: where it actually lands
Reflection published a wide benchmark table, and to its credit, the table includes models that beat Beam. Below are the rows I'd look at for agent and coding work, and every number comes from Reflection's own table. "NR" stands for not reported.
| Benchmark | Beam | GLM-5.2 | Kimi K3 | Qwen3.8 Max | DeepSeek V4.1 Flash | Inkling |
|---|---|---|---|---|---|---|
| SWE-bench Verified | 80.9 | NR | NR | NR | NR | 77.6 |
| SWE-bench Pro v1 | 65.5 | 62.1 | NR | 67.7 | NR | 54.3 |
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.3 | 86.6 | 90.6 | 63.8 |
| DeepSWE v1.1 | 44.4 | 44.0 | 68.0 | 51.0 | 74.2 | NR |
| MCP Atlas | 78.7 | 77.8 | 82.3 | 84.5 | NR | 76.0 |
| tau3 banking | 38.0 | 37.1 | 37.1 | 55.2 | NR | 25.0 |
| HLE (no tools) | 36.2 | 40.5 | 46.9 | 43.6 | 39.1 | 29.7 |
| GPQA Diamond | 90.5 | 91.2 | 93.5 | 92.6 | 90.9 | 87.2 |
When you read across the rows, the picture is fairly clear: Beam trades blows with GLM-5.2 and comfortably beats Inkling, but trails Kimi K3, Qwen3.8 Max and DeepSeek V4.1 Flash wherever those were reported. Reflection's own framing lines up with that, saying Beam is "competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks," while "frontier open models like Kimi K3 remain ahead on raw capability."
The headline 80.9 on SWE-bench Verified needs a footnote, and an X user named LuChai wrote it well:
"The 80.9 on SWE Bench Verified is the standout, but worth flagging that GLM-5.2 and Qwen 3.8 Max are both NR there"
There's also no independent referee so far. As of October 8, Beam isn't listed on Artificial Analysis (its 696-model list has no Beam entry), and it's missing from Sierra's tau-bench leaderboard too. For context, Artificial Analysis currently scores Qwen3.8 Max at 45.4, Kimi K3 at 43.6 and GLM-5.2 at 33.7 on its Intelligence Index. If Beam really does sit near GLM-5.2, I'd expect it somewhere in the 30s once it gets tested, and Mistral Large 4, another recent Western release, makes a useful yardstick for that.
One more thing I liked is that Reflection shows its RL run kept improving without a plateau, which is a decent sign for the next model in the series.

What will Reflection AI Beam cost?
Short answer: nobody outside the waitlist knows yet. Both the Reflection platform and its /pricing page show nothing but a sign-in screen. According to the commercial terms, fees are "set forth in the pricing documentation" on the platform, and that documentation is behind the same login.
The error and billing docs do reveal one thing, which is the shape of the billing:
- A credit or prepaid balance (
402 insufficient_credits,prepaid_balance_exhausted) - A verified card before some organizations can create API keys (
403 payment_method_required) - Billing tiers and a daily token allowance set by your plan
Since there's no Beam price to quote, here's the range Beam would have to land in for its efficiency story to matter. Below are the first-party API rates for the models in Reflection's comparison table, per 1M tokens:
| Model | Input | Cached input | Output | Source |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 | Kimi platform |
| Qwen3.8 Max | $2.00 | $0.25 | $6.00 | Qwen Cloud |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | Z.ai pricing |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 (peak) | DeepSeek pricing |
| Reflection Beam | Not published | Not published | Not published | Waitlist only |
Say Beam lands near GLM-5.2 on quality. Then the efficiency argument only pays off if it's priced at or under GLM-5.2's $1.40 / $4.40, and it still has to explain itself against DeepSeek V4.1 Flash, which scores higher on most of Reflection's own rows at a fraction of those rates. For how those numbers play out per task, see my deeper breakdowns of GLM pricing and Qwen3.8 Max pricing.
Self-hosting changes the math, though not the hardware bill. One X reply to a "run it on your desk" post made the counterpoint in blunt terms:
"m5 studio 512GB is like $20 grand... this is realistically a $40k investment, to use 6 month old AI."
Data terms to read before you send real data
Reflection's terms split the service into two kinds. On paid Standard Services, you license your content to Reflection "solely to operate and provide the Services." On free or Discounted Services the license is much broader, and it includes the right to "create derivative works of, and otherwise exploit" your content. In both cases content is kept for 30 days for safety and compliance, with flagged content kept for up to 12 months. If you're testing Beam with customer conversations on a discounted early-access plan, that clause is the one I'd read twice.
How to use Beam right now
Once you're off the waitlist, the API is easy to drop in. Reflection runs an OpenAI-compatible endpoint at https://api.reflection.ai/openai/v1, which means code built on the OpenAI SDKs works after you swap the base URL and the key. The model ID is Beam-501B-A23B.
There are a few gotchas in the docs that are worth knowing before you build:
- Only Chat Completions and Models are supported. No Responses API, embeddings, batch, files or assistants.
stopsequences are accepted but ignored. Generation won't stop at them, so don't rely on them for parsing.- Reasoning can't be turned off. Pick
lowfor speed,maxfor hard tasks. Reasoning text comes back in a separatereasoning_contentfield. nis fixed at 1 and logprobs are off. If your eval harness samples multiple completions per call, it needs a loop.- Tool calling and structured outputs work, including
parallel_tool_callsandjson_schema.
For coding work, Reflection ships its own terminal agent, Mirror CLI, with Beam as the default model. It's still in early beta, and it supports parallel worktrees and AGENTS.md instructions, plus MCP connections to tools like Linear and Notion. Reflection's quickstart also covers how to wire Beam into Pi, OpenCode and Hermes Agent, which is useful if you already live in one of those.
If you're still choosing a harness, that side is covered in my Pi coding agent and agentic coding CLI posts.

Reflection's own demos lean hard into this agent angle. My favorite one is the live NYC subway map, where Beam searched the MTA docs and worked out auth, then found the subway geometry and built both the frontend and the backend:

One demo did need a correction after launch. Reflection's "Land or Water" demo first claimed the puzzle was only days old, so it couldn't be in the training data. After an HN user pointed out it was older, a Reflection team member replied that the caption had been changed. It's a small thing and it was handled openly, but it's worth keeping in mind when you read the demos.
What developers are saying about Beam
I went through the Hacker News thread and r/LocalLLaMA, along with X and LinkedIn, and the same three themes kept coming up.
"Show me the weights." Until the weights ship, a lot of people simply won't count Beam as open. The moderators of r/LocalLLaMA removed the launch thread for exactly that reason:
"Until such time that weights become available, this is off-topic for LocalLLaMA."
"It's a generation behind, but it's American." Most commenters put Beam somewhere around GLM-5.2, which is a June model. The excitement is more about where it was built, and for teams that can't use Chinese-origin models, that is the whole decision:
"I'm in a similar situation to you where we can't use Chinese AI models but we want to host on-prem due to data sovereignty concerns. ... I'm also keeping an eye on Reflection Beam"
"Nice efficiency claim, now let someone check it." Even the positive takes come with that caveat attached. Tyler Folkman, Chief AI Officer at JobNimbus, summed up how practitioners are going to size it:
"Nobody outside Reflection has verified any of this yet. ... When Beam's weights ship, I'm sizing them for my boxes by active parameters and quantized footprint, not the leaderboard."
Reflection also got some credit for candor. One HN commenter praised Reflection for admitting Kimi K3 is ahead "instead of making false claims." I agree with that, since a launch post that names the models beating you is rarer than it should be.
Is Beam a good model for customer support?
This is the question I care about the most, since my days go into the AI that answers support tickets. The good news is that Reflection reported a score on a benchmark that really looks like support work.
tau3 banking comes from Sierra's tau-bench. It drops the model into a simulated bank help desk with 698 policy documents and asks it to resolve customer requests through tools. A task passes only when the customer's account ends up in the right state, which means a polite but wrong answer scores zero. Each task needs about 18.6 documents and 9.5 tool calls.

Beam's 38.0 is roughly level with Kimi K3 and GLM-5.2 (both 37.1 on Sierra's leaderboard) and well behind Qwen3.8 Max at 55.2. Reflection doesn't say which test setup gave its 38.0, so I'd treat it as close rather than exact. The whole field also drops once you ask for consistency: on Sierra's board, Qwen3.8 Max falls from 55.2 to 35.1 when a task has to pass all four tries.
That last number is the real lesson for support teams, and I covered it in depth in my breakdown of the best AI model for support tickets: even the best model gets a policy-heavy ticket right every time only about a third of the time. Swapping in a new model doesn't fix that.
What fixes it is the layer around the model, which means grounding answers in your real knowledge base and past tickets and escalating low-confidence replies to a human, plus testing on your own history before anything goes live. My implementation guide walks through that rollout order.
I've seen that failure play out on real queues. A vehicle telematics team on Zendesk had a bot confidently telling customers "yes, we support your car model" for brands that weren't in their database, all because one help article said "we support all models." No benchmark score would have caught that, but a simulation run against their own past tickets would have.
So for support specifically, Beam is a reasonable engine candidate once it has a price and the weights are out, more so if your company rules out Chinese-origin models. By itself it isn't a support agent, though. For that part you need an AI helpdesk agent that knows your queue, and my AI helpdesk roundup compares the ones that do.
Who is Reflection AI?
Reflection was co-founded in 2024 by Misha Laskin (CEO) and Ioannis Antonoglou, both former Google DeepMind researchers. In his launch post, Laskin traces his own switch from physics to AI back to AlphaGo in 2016.
There's a lot of money behind it. Reflection's own post announced a $2 billion raise in 2025 from investors including NVIDIA, Sequoia, Lightspeed, DST, CRV and Eric Schmidt, and its newsroom lists a $25 billion pre-money valuation for its April 2026 round. In the launch post, Laskin also makes the case for why they went open:
"Open also means safer. The best defence is a distributed one."
Antonoglou added on X that the team is "already training the next, larger model." Beam, in other words, is explicitly the first in a series.
One clarification came up on HN as well: this Reflection has no relation to the 2024 "Reflection 70B" model. A commenter corrected the mix-up in the launch thread, and the person who had raised it took it back.
So, should you wait for Beam?
Here's where I'd land, depending on the reader:
| If you are... | My call |
|---|---|
| A team that needs a US-built open model (gov, regulated, "no Chinese-origin" policy) | Join the waitlist now and plan an eval for early November |
| Building coding agents on a budget today | Use DeepSeek V4 Flash or GLM now, re-test Beam once it's on Artificial Analysis |
| Chasing the strongest open model | Kimi K3 or Qwen3.8 Max, per Reflection's own table |
| Planning to self-host | Wait for the FP8 and NVFP4 weights, then size by memory for 501B, not compute for 23B |
| Picking a model for a support queue | Pick the agent first, then the model; see below |
eesel for support teams watching the model race
If you landed here because you're picking a model to answer customers, the shortcut is that you don't have to bet your queue on any single launch. eesel's AI helpdesk teammate plugs into Zendesk, Freshdesk or Gorgias and learns from your past tickets and help center, then runs a simulation on your real ticket history before it replies to anyone. When a better or cheaper model ships, the teammate gets better and you don't have to rebuild anything.

It also works for the terminal crowd that reads a Beam post. With the eesel CLI you can run the same teammate from a shell, for things like connecting a helpdesk, editing its standing instructions, approving or denying pending actions and reading its activity log. Every command returns JSON, and a --dry-run flag shows the exact call before it's sent. On top of that, every workspace is an MCP server, so a coding agent like Claude Code can drive it directly.
Results tend to come quickly when the agent is grounded in your own data. Here's how Gridwise put it after their trial:
"In the first month, eesel is resolving 73% of our tier 1 requests."
Try eesel free with 100 credits, no card needed, or see the paid plans on the pricing page.
Frequently Asked Questions
What is Reflection AI Beam?
How much does Reflection AI Beam cost?
Can I download the Reflection AI Beam weights yet?
Is Reflection AI Beam better than Kimi K3 or Qwen3.8 Max?
What context window does Reflection AI Beam have?
Can Reflection AI Beam work as a customer support agent?
Does Reflection AI Beam train on my API data?

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.







