
What I mean by a review, and what I could not test
Muse Spark 1.1 shipped on 9 July 2026, so it has had about four weeks in the wild rather than four days. That matters. There is independent data now, not just launch-day vibes. The wider company picture sits in my Meta AI overview.
There is less first-hand reporting than a frontier launch usually generates, and the reason for that is structural. The Meta Model API opened US-only. Developers reported being blocked on day one from Vietnam, Canada and Argentina, and it was not on OpenRouter until about a week later. One commenter put the practical effect plainly:
"This not being available on Openrouter really makes it hard to test. I was going to compare vs Grok 4.5 and GPT-5.6 Luna, but I don't want to deal with signing up for Meta for it unless it checks out."
So a 413-point launch thread produced only a handful of real usage reports. Everything below leans on three source types: Artificial Analysis' independent runs and Meta's own 105-page evaluation report, plus the small number of engineers who actually got the thing working. Where nobody has tested something, I say so instead of filling the gap.
The one chart that explains every other result
Here is Muse Spark 1.1's rank across the benchmarks that make up the Intelligence Index. Same model, same week, one lab running all of the tests.

Most models are boring in a useful way. They sit in roughly the same percentile wherever you test them, so a single number predicts the rest. This one does not. The spread between its best and worst rank is wider than the entire gap between the top and bottom of most price tiers.
That is why the headline framing everyone repeats, 50.6 against Opus 5's 60.7, tells you almost nothing you can act on. It averages one excellent result with one poor one. What you need to know is which of the two you are buying.
Where it actually wins
Three things here are real, and I would rather the criticism further down did not bury them.
It writes code well. Not "well for the price". Well. On SciCode it beats Claude Opus 5, GPT-5.6 Sol and Grok 4.5. The "Meta is behind on coding" line that followed the launch is about long agent runs rather than about generating code, and those are two different jobs.
| Model | SciCode |
|---|---|
| Kimi K3 (max) | 58.7% |
| Muse Spark 1.1 (xhigh) | 58.2% |
| GPT-5.6 Sol (high) | 56.9% |
| Claude Opus 5 (max) | 55.7% |
| Grok 4.5 (high) | 54.1% |
| Gemini 3.6 Flash | 52.7% |
| DeepSeek V4 Flash | 49.9% |
It is the fastest thing on the board. 217.4 output tokens per second against 55.7 for Claude Opus 5, and a little ahead of Gemini 3.6 Flash at 213.5. For anything user-facing, users feel a 4x throughput difference before anyone measures it.
The cached rate is the best-designed part of the product. At $0.15 per million that is an 88% discount, applied automatically, with no cache key for you to manage. One engineer flagged why the ratio matters more than the sticker price:
"The cached input pricing is a good ratio. Compare with Grok 4.5 which came out at $2/$6 but then quietly charges $0.50 per 1M cached input tokens. That's as high as Opus 4.8!"
Put those together and the good use case takes an obvious shape: short prompts, a stable system prompt, huge volume, one decision per call. Batch classification, tagging, extraction, routing. The full rate card lives in my Muse Spark 1.1 pricing breakdown. For the rest of the field, start with Claude pricing. The OpenAI lineup sits in one table in every OpenAI model.
Where it comes apart
Meta positioned this as an agent model. Its announcement leads on tool use and multi-step workflows, the line that separates agents from chatbots. So the fair test is the independent agentic evals. This is where the review turns.
| Agentic benchmark | Muse Spark 1.1 | Best in set | Where it lands |
|---|---|---|---|
| AA-Briefcase Elo | 868.9 | Claude Opus 5 at 1718.7 | 18th of 19 |
| GDPval-AA v2 Elo | 1370.6 | Claude Opus 5 at 1852.0 | Last of 7 |
| Tau-3 Banking | 25.2% | Kimi K3 at 34.0% | 9th of 10 |
The contradiction there is hard to miss. The model sold on agentic tool use is the weakest agentic performer on the independent board. On GDPval it sits 481 Elo behind Opus 5, and the confidence intervals do not overlap anything above it. On agentic tool use it trails DeepSeek V4 Flash, which costs $0.14 in and $0.28 out.
Meta does not entirely dispute this. Its own evaluation report concedes that "for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models." The admission is measured against Opus 4.8 and GPT-5.5, both a generation behind what you would buy today.
The community read has settled in roughly the same place. One commenter, watching it appear near the top of a coding leaderboard, took that as evidence the leaderboard was broken:
"your benchmark is obviously flawed when the top 3 models for Typescript (Combined) are Grok 4.5, Muse Spark 1.1 (lol), Gemini 3.5! Flash"
The sharpest framing came out of a four-model build-off. Someone there was confused that Muse Spark produced the best single artefact in the set and still scored 2 out of 5. The clarification is the whole model in one line:
"2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures"
High ceiling. Low floor. Everything else in this review follows from that.
The longest single sitting anyone has published lands in the same spot. Three hours with the model, from a developer running it through a coding CLI:
"I just tried it for 3 hours straight, and I have to say, I'm disappointed. I don't know what has gotten into Meta these days, but gotta remember that they were the Llama creator that sparked the OSS of models. Even closed source (at least now) I would say it's between Kimi 2.7 and GLM 5.2, not even close to Opus 4.8 medium/Sonnet5"
Hold that one loosely. It is a single person's afternoon, and he reports no measured numbers. Almost nobody does. Across four weeks of Reddit and Hacker News there are roughly a dozen real first-hand reports of running this model, and not one of them publishes a figure they measured themselves. Nobody has stress-tested the 1M window in public either. Partly the region lock, partly that a model you cannot get an API key for does not get benchmarked by hobbyists.
Should you hand this job to Muse Spark 1.1?
Pick the job you actually have. The answer changes more than the price does.
This is the job it was built for. Short input, one decision, huge volume, and the automatic cached rate does the rest.
It runs at 217.4 output tokens per second, the fastest on the board, and a stable system prompt bills at $0.15 per million instead of $1.25.
The coding weakness everyone repeats is about long agent runs, not about writing code. Asked for a single function, it is near the top of the field.
On SciCode it places 2nd of 8 at 58.2%, above Claude Opus 5 at 55.7% and every GPT-5.6 variant tested.
Long-horizon work is where it drifts, and the retries quietly hand the savings back. Keep a stronger model on the multi-hour runs.
AA-Briefcase puts it 18th of 19 at 868.9 Elo, below Gemini 3.6 Flash and below Nemotron 3 Ultra.
The million-token window is the headline feature and one of the weakest measured results in the model. Retrieve first, then prompt. Do not dump the archive in.
Long-context reasoning lands 25th of 26, and Meta's own report scores 54.1 on MRCR v2 against GPT-5.5's 74.0.
The computer-use demo is real and checkable rather than staged, and it decides per step whether to script or click. It still does not lead the benchmark it markets hardest.
Meta's own report puts it at 80.8 on OSWorld-Verified, behind Claude Opus 4.8 at 83.4.
Not inventing facts is not the same as knowing when to stay quiet, and this is the model's weakest alignment dimension by Meta's own testing.
Meta's Petri run scores input hallucination at 1.67, worse than GPT-5.5 at 1.22 and Claude Opus 4.8 at 1.34.
The failure mode Meta documents itself
This is the part I would want to know before wiring it into anything. It is buried in the docs, not the launch post.
Muse Spark 1.1 keeps its chain of thought private, which is normal now. The unusual part is what happens on the most common integration path. Meta's own reasoning documentation states that on Chat Completions, the reasoning_content field is redacted to empty before the response reaches the caller, "so there is nothing to replay and each turn reasons from scratch."
Meta's coding agents guide spells the consequence out: the model "can lose the thread of its own prior thinking and behave erratically: repeating work it already did, contradicting earlier steps."
A model sold on multi-step agentic work forgets its own reasoning between turns on the path most tools use by default. Multi-turn reasoning survives only on the Responses API, where the server keeps the context for you through previous_response_id.
None of that is theoretical. The best hands-on report from the launch describes exactly this class of problem, and it comes from an engineer who got Codex running against the API inside a container:
"It's some kind of parsing or integration error due to what I think is codex not anticipating server-side tool calling and how meta treats those ids... first couple times running codex with muse, it would fail on its first non-web search call."
He fixed it, and he stayed positive about the model. The point is that two independent surfaces, Meta's own docs and the first person to wire it into a real harness, land on the same root cause. The agentic plumbing is bespoke, and third-party harnesses trip over it.
The hallucination number everyone quotes is the wrong one
Here I have to correct a flattering reading that has been going around, my own earlier coverage included.
Muse Spark 1.1 does have a good non-hallucination rate, meaning it declines to answer instead of inventing things reasonably often. True, and worth something. But the composite knowledge reliability index, which rewards correct answers and penalises hallucinations, puts it at 18.0. That is the lowest of the twelve configurations listed on its own model page. Claude Opus 5 scores 31.3.
Meta's own alignment testing agrees, in blunter language than any competitor published. From the Petri 3.0 assessment in its evaluation report: "input hallucination (1.67) is the primary one, higher than GPT-5.5 (1.22) and Claude 4.8 Opus (1.34)." The same passage flags elevated deception toward users, and elevated overrefusal.
So two independent sources, one of them Meta itself, land on the same conclusion: hallucination is this model's weakest dimension, not its strongest.
Even that is not the number I would decide a support rollout on, though. This is:

Every hallucination benchmark measures whether a model invents facts about the world. Almost no support failure looks like that. A B2B technical support team we worked with, running roughly 200 tickets a month on Zendesk and scaling toward 2,000, hit the real version of the problem. Their bot told customers it supported vehicles that were not in their database, because their own help centre said "we support all models." The model was faithful to its source. The source was wrong.
No score on any leaderboard catches that. It is why training AI on a knowledge base is a content problem before it is a model problem, and why confidence thresholds do more for answer quality than a reasoning upgrade will.
The same logic applies upstream. Getting ticket triage right moves resolution rate more reliably than a model swap does. The failure patterns are consistent enough that AI chatbot problems catalogues them better than any model card.
Speed is real, but read the right clock
The throughput number is real, and it is the model's best single stat. The latency number needs a correction.
| Measure | Muse Spark 1.1 | Claude Opus 5 (max) |
|---|---|---|
| Output speed | 217.4 tok/s | 55.7 tok/s |
| Time to first token | 2.89s | - |
| Thinking time before first answer | 9.20s | - |
| Time to first answer token | 12.09s | 51.22s |
Meta's raw 2.89s time to first token is the best on the board. It is also not what a user experiences. Add 9.2 seconds of hidden reasoning and the first useful token arrives at 12.09s. Still comfortably better than Opus 5, so the conclusion holds. Four times slower than the headline implies, though.
That same hidden reasoning drives the cost story. 68% of billed output tokens are thinking the caller never sees, at 15,164 reasoning tokens against 7,232 answer tokens per task. On Artificial Analysis' full index run, $360 of the roughly $548 total went on reasoning. Cheap tokens, expensive thinking habit. It is the reason AI customer service cost never tracks the rate card.
Four footguns worth knowing before you start
None of these are dealbreakers. Each one costs an afternoon if you meet it cold.
tool_choiceonly accepts"auto". You cannot force a specific tool, and there is no"required"or"none". Anything else comes back as a 400.- Structured output defaults to off. Under
strict: false, Meta's tool-calling reference warns that generated arguments "are not guaranteed to validate against" your schema. Validate before you execute anything. - Claude Code needs three separate changes. A base URL with no
/v1, all five model aliases repointed tomuse-spark-1.1, thenANTHROPIC_AUTH_TOKENin place ofANTHROPIC_API_KEY. - MCP is claimed but not documented. The launch post says the model generalises to MCP servers and custom skills. The tool-calling reference documents neither of them. Developer-defined tools are the only extensibility path actually shown.
There is also no server-side cap on runaway custom-tool loops. max_tool_calls limits Meta's built-in tools only, worth knowing if you have ever watched an AI agent loop spin. The delegation half of this is covered in subagent orchestration, and the case for testing it yourself sits in agent evals.
What Meta still has not published
Gaps are review material too, and this list runs longer than it should four weeks after launch.
| Not published | Why it matters |
|---|---|
| Max output tokens | You cannot size a request ceiling |
| Knowledge cutoff | You cannot reason about staleness |
| Parameter count | No architecture detail of any kind |
| Regional allowlist | Everything known comes from blocked users |
| Data retention duration | Paid prompts skip training, retention is unstated |
| Uptime or availability SLA | Nothing at all |
| Minimum cacheable prefix | The $0.15 rate leans on an undocumented variable |
Meta also does not appear on the official Terminal-Bench leaderboard, so the self-reported score there has no third-party confirmation. My Muse Spark 1.1 overview covers the methodology dispute around that number. It is still unresolved.
Who should buy it, and who should skip

Buy it if you run high-volume, short-horizon work: classification, tagging, extraction, routing, single-file code generation, anything where one call produces one answer and you make millions of those calls. The speed is real and the cached rate is excellent, and at roughly $0.29 per task the savings against Opus 5 pricing are big enough to matter.
Skip it if your workload is long-running agents, deep research runs, or anything that has to reason across a large document set. The independent agentic numbers are not close. This is the one case where paying for GPT-5.6 or Opus 5 comes out cheaper once you count the retries. For coding harnesses specifically, Codex is the better-documented path today.
Wait if you are outside the US, or you need a published retention policy, or you need open weights. On that last one, Kimi K3 is the closest near-frontier option with weights you can actually hold. The closed-weights turn is the part of this launch the community has forgiven least.
The moderate read, from someone tracking the leaderboard, is fair enough that I would sign it:
"New respect for Meta Muse Spark. It seems to sit at a lot of sweet spots in the leader board. It's not the best at anything in particular, but it balances cost and performance quite well."
What this changes for a support queue
Almost nothing. I say that as someone whose job is partly to notice when a new model does change something.
Every few weeks a cheaper, faster model ships and somebody asks whether it rewrites the plan for their helpdesk. The honest answer is that the model was never the constraint. In eesel's own cross-validated trials, when agents rewrote an AI draft, roughly 65% of the edits were length and tone. About 20% needed data the AI could not reach in an ERP or logistics system. Only around 5% were the AI being factually wrong. A better model addresses that last 5%. The rest is prompt engineering, retrieval, agent coaching on your team's own sent replies, and integration depth.
Which is also why copilot-style drafting is where most teams should start, the pattern behind agent assist tools, and why a clean escalation path matters more than a leaderboard position does. If you are building the business case rather than the stack, AI vs human cost is the more useful frame.
The failure I watch for hardest in production is the one no benchmark on Meta's launch page measures. An agent narrating a search it never ran. Reporting files it never saved. An agent that claims to have done the work is a harder problem than an agent doing the work badly, and a model that forgets its own reasoning between turns is not the one I would trust to self-report.
Try eesel for support, not a raw model key
If what you actually want is an AI answering customer tickets, the useful question has little to do with which model tops the index this month. It is whether you can prove the thing is safe before it replies to anyone.
That is the part eesel is built around. You can run an AI agent in simulation against your own historical tickets, and read its answers on real past conversations before a single customer sees one. If the answers are not there yet, you start in copilot mode, where it drafts and your team still sends.
Setup is a helpdesk connection, not a project. It plugs into Zendesk in a few minutes, reads the help centre you already wrote, and bills per resolved ticket instead of per million tokens, so the invoice tracks work done rather than how verbose the model felt that day. Free to try, and the pricing is public.
The verdict
Muse Spark 1.1 is a good model wearing the wrong label. Meta sold an agent, and the agent benchmarks are its worst results. On the evidence it is the best fast-and-cheap single-shot model available right now, with a strong code-generation score and a cached rate nobody else matches.
Judge it on the job you have rather than the category Meta filed it under, and it becomes easy to place. Buy it for the millions of short calls. Keep something stronger on the long runs. And do not let a 1M-token window talk you out of retrieval. At roughly an eighth of Opus 5's cost per task, being second-best at the right jobs is a perfectly good business.
Frequently Asked Questions
Is Meta Muse Spark 1.1 any good?
How does Muse Spark 1.1 compare to Claude Opus 5?
Is Muse Spark 1.1 fast?
Does Muse Spark 1.1 hallucinate?
Can I use Muse Spark 1.1 with Claude Code or Codex?
/v1, all five model aliases repointed, and ANTHROPIC_AUTH_TOKEN rather than an API key. The bigger trap is that the Chat Completions path drops the model's reasoning between turns, so it repeats and contradicts itself on long runs. Use the Responses API instead.Is Muse Spark 1.1 worth it for customer support?
Where can I use the Meta Model API?

Article by
Kurnia Kharisma Agung Samiadjie
Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.








