Meta Muse Spark 1.1 review: a high ceiling and a low floor

Kurnia Kharisma Agung Samiadjie
Written by

Kurnia Kharisma Agung Samiadjie

Katelin Teen
Reviewed by

Katelin Teen

Last edited August 5, 2026

Expert Verified
A developer workbench where one AI model finishes one task cleanly and stalls on the one beside it, in Meta's blue brand colour

What I mean by a review, and what I could not test

Muse Spark 1.1 shipped on 9 July 2026, so it has had about four weeks in the wild rather than four days. That matters. There is independent data now, not just launch-day vibes. The wider company picture sits in my Meta AI overview.

There is less first-hand reporting than a frontier launch usually generates, and the reason for that is structural. The Meta Model API opened US-only. Developers reported being blocked on day one from Vietnam, Canada and Argentina, and it was not on OpenRouter until about a week later. One commenter put the practical effect plainly:

Hacker News

"This not being available on Openrouter really makes it hard to test. I was going to compare vs Grok 4.5 and GPT-5.6 Luna, but I don't want to deal with signing up for Meta for it unless it checks out."

So a 413-point launch thread produced only a handful of real usage reports. Everything below leans on three source types: Artificial Analysis' independent runs and Meta's own 105-page evaluation report, plus the small number of engineers who actually got the thing working. Where nobody has tested something, I say so instead of filling the gap.

The one chart that explains every other result

Here is Muse Spark 1.1's rank across the benchmarks that make up the Intelligence Index. Same model, same week, one lab running all of the tests.

One model ranked across six benchmarks, from 3rd of 26 on writing code down to 25th of 26 on long-context reasoning, with a bracket spanning the whole distance labelled same model
One model ranked across six benchmarks, from 3rd of 26 on writing code down to 25th of 26 on long-context reasoning, with a bracket spanning the whole distance labelled same model

Most models are boring in a useful way. They sit in roughly the same percentile wherever you test them, so a single number predicts the rest. This one does not. The spread between its best and worst rank is wider than the entire gap between the top and bottom of most price tiers.

That is why the headline framing everyone repeats, 50.6 against Opus 5's 60.7, tells you almost nothing you can act on. It averages one excellent result with one poor one. What you need to know is which of the two you are buying.

Where it actually wins

Three things here are real, and I would rather the criticism further down did not bury them.

It writes code well. Not "well for the price". Well. On SciCode it beats Claude Opus 5, GPT-5.6 Sol and Grok 4.5. The "Meta is behind on coding" line that followed the launch is about long agent runs rather than about generating code, and those are two different jobs.

ModelSciCode
Kimi K3 (max)58.7%
Muse Spark 1.1 (xhigh)58.2%
GPT-5.6 Sol (high)56.9%
Claude Opus 5 (max)55.7%
Grok 4.5 (high)54.1%
Gemini 3.6 Flash52.7%
DeepSeek V4 Flash49.9%

It is the fastest thing on the board. 217.4 output tokens per second against 55.7 for Claude Opus 5, and a little ahead of Gemini 3.6 Flash at 213.5. For anything user-facing, users feel a 4x throughput difference before anyone measures it.

The cached rate is the best-designed part of the product. At $0.15 per million that is an 88% discount, applied automatically, with no cache key for you to manage. One engineer flagged why the ratio matters more than the sticker price:

Hacker News

"The cached input pricing is a good ratio. Compare with Grok 4.5 which came out at $2/$6 but then quietly charges $0.50 per 1M cached input tokens. That's as high as Opus 4.8!"

Put those together and the good use case takes an obvious shape: short prompts, a stable system prompt, huge volume, one decision per call. Batch classification, tagging, extraction, routing. The full rate card lives in my Muse Spark 1.1 pricing breakdown. For the rest of the field, start with Claude pricing. The OpenAI lineup sits in one table in every OpenAI model.

Where it comes apart

Meta positioned this as an agent model. Its announcement leads on tool use and multi-step workflows, the line that separates agents from chatbots. So the fair test is the independent agentic evals. This is where the review turns.

Agentic benchmarkMuse Spark 1.1Best in setWhere it lands
AA-Briefcase Elo868.9Claude Opus 5 at 1718.718th of 19
GDPval-AA v2 Elo1370.6Claude Opus 5 at 1852.0Last of 7
Tau-3 Banking25.2%Kimi K3 at 34.0%9th of 10

The contradiction there is hard to miss. The model sold on agentic tool use is the weakest agentic performer on the independent board. On GDPval it sits 481 Elo behind Opus 5, and the confidence intervals do not overlap anything above it. On agentic tool use it trails DeepSeek V4 Flash, which costs $0.14 in and $0.28 out.

Meta does not entirely dispute this. Its own evaluation report concedes that "for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models." The admission is measured against Opus 4.8 and GPT-5.5, both a generation behind what you would buy today.

The community read has settled in roughly the same place. One commenter, watching it appear near the top of a coding leaderboard, took that as evidence the leaderboard was broken:

Hacker News

"your benchmark is obviously flawed when the top 3 models for Typescript (Combined) are Grok 4.5, Muse Spark 1.1 (lol), Gemini 3.5! Flash"

The sharpest framing came out of a four-model build-off. Someone there was confused that Muse Spark produced the best single artefact in the set and still scored 2 out of 5. The clarification is the whole model in one line:

Hacker News

"2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures"

High ceiling. Low floor. Everything else in this review follows from that.

The longest single sitting anyone has published lands in the same spot. Three hours with the model, from a developer running it through a coding CLI:

Reddit

"I just tried it for 3 hours straight, and I have to say, I'm disappointed. I don't know what has gotten into Meta these days, but gotta remember that they were the Llama creator that sparked the OSS of models. Even closed source (at least now) I would say it's between Kimi 2.7 and GLM 5.2, not even close to Opus 4.8 medium/Sonnet5"

Hold that one loosely. It is a single person's afternoon, and he reports no measured numbers. Almost nobody does. Across four weeks of Reddit and Hacker News there are roughly a dozen real first-hand reports of running this model, and not one of them publishes a figure they measured themselves. Nobody has stress-tested the 1M window in public either. Partly the region lock, partly that a model you cannot get an API key for does not get benchmarked by hobbyists.

Should you hand this job to Muse Spark 1.1?

Pick the job you actually have. The answer changes more than the price does.

Buy

This is the job it was built for. Short input, one decision, huge volume, and the automatic cached rate does the rest.

It runs at 217.4 output tokens per second, the fastest on the board, and a stable system prompt bills at $0.15 per million instead of $1.25.

Buy

The coding weakness everyone repeats is about long agent runs, not about writing code. Asked for a single function, it is near the top of the field.

On SciCode it places 2nd of 8 at 58.2%, above Claude Opus 5 at 55.7% and every GPT-5.6 variant tested.

Skip

Long-horizon work is where it drifts, and the retries quietly hand the savings back. Keep a stronger model on the multi-hour runs.

AA-Briefcase puts it 18th of 19 at 868.9 Elo, below Gemini 3.6 Flash and below Nemotron 3 Ultra.

Skip

The million-token window is the headline feature and one of the weakest measured results in the model. Retrieve first, then prompt. Do not dump the archive in.

Long-context reasoning lands 25th of 26, and Meta's own report scores 54.1 on MRCR v2 against GPT-5.5's 74.0.

Worth a look

The computer-use demo is real and checkable rather than staged, and it decides per step whether to script or click. It still does not lead the benchmark it markets hardest.

Meta's own report puts it at 80.8 on OSWorld-Verified, behind Claude Opus 4.8 at 83.4.

Careful

Not inventing facts is not the same as knowing when to stay quiet, and this is the model's weakest alignment dimension by Meta's own testing.

Meta's Petri run scores input hallucination at 1.67, worse than GPT-5.5 at 1.22 and Claude Opus 4.8 at 1.34.

The failure mode Meta documents itself

This is the part I would want to know before wiring it into anything. It is buried in the docs, not the launch post.

Muse Spark 1.1 keeps its chain of thought private, which is normal now. The unusual part is what happens on the most common integration path. Meta's own reasoning documentation states that on Chat Completions, the reasoning_content field is redacted to empty before the response reaches the caller, "so there is nothing to replay and each turn reasons from scratch."

Meta's coding agents guide spells the consequence out: the model "can lose the thread of its own prior thinking and behave erratically: repeating work it already did, contradicting earlier steps."

A model sold on multi-step agentic work forgets its own reasoning between turns on the path most tools use by default. Multi-turn reasoning survives only on the Responses API, where the server keeps the context for you through previous_response_id.

None of that is theoretical. The best hands-on report from the launch describes exactly this class of problem, and it comes from an engineer who got Codex running against the API inside a container:

Hacker News

"It's some kind of parsing or integration error due to what I think is codex not anticipating server-side tool calling and how meta treats those ids... first couple times running codex with muse, it would fail on its first non-web search call."

He fixed it, and he stayed positive about the model. The point is that two independent surfaces, Meta's own docs and the first person to wire it into a real harness, land on the same root cause. The agentic plumbing is bespoke, and third-party harnesses trip over it.

The hallucination number everyone quotes is the wrong one

Here I have to correct a flattering reading that has been going around, my own earlier coverage included.

Muse Spark 1.1 does have a good non-hallucination rate, meaning it declines to answer instead of inventing things reasonably often. True, and worth something. But the composite knowledge reliability index, which rewards correct answers and penalises hallucinations, puts it at 18.0. That is the lowest of the twelve configurations listed on its own model page. Claude Opus 5 scores 31.3.

Meta's own alignment testing agrees, in blunter language than any competitor published. From the Petri 3.0 assessment in its evaluation report: "input hallucination (1.67) is the primary one, higher than GPT-5.5 (1.22) and Claude 4.8 Opus (1.34)." The same passage flags elevated deception toward users, and elevated overrefusal.

So two independent sources, one of them Meta itself, land on the same conclusion: hallucination is this model's weakest dimension, not its strongest.

Even that is not the number I would decide a support rollout on, though. This is:

Two panels contrasting a benchmark question the model correctly declines with a help desk question where the knowledge base overpromises and the model confidently agrees, both scored as correct
Two panels contrasting a benchmark question the model correctly declines with a help desk question where the knowledge base overpromises and the model confidently agrees, both scored as correct

Every hallucination benchmark measures whether a model invents facts about the world. Almost no support failure looks like that. A B2B technical support team we worked with, running roughly 200 tickets a month on Zendesk and scaling toward 2,000, hit the real version of the problem. Their bot told customers it supported vehicles that were not in their database, because their own help centre said "we support all models." The model was faithful to its source. The source was wrong.

No score on any leaderboard catches that. It is why training AI on a knowledge base is a content problem before it is a model problem, and why confidence thresholds do more for answer quality than a reasoning upgrade will.

The same logic applies upstream. Getting ticket triage right moves resolution rate more reliably than a model swap does. The failure patterns are consistent enough that AI chatbot problems catalogues them better than any model card.

Speed is real, but read the right clock

The throughput number is real, and it is the model's best single stat. The latency number needs a correction.

MeasureMuse Spark 1.1Claude Opus 5 (max)
Output speed217.4 tok/s55.7 tok/s
Time to first token2.89s-
Thinking time before first answer9.20s-
Time to first answer token12.09s51.22s

Meta's raw 2.89s time to first token is the best on the board. It is also not what a user experiences. Add 9.2 seconds of hidden reasoning and the first useful token arrives at 12.09s. Still comfortably better than Opus 5, so the conclusion holds. Four times slower than the headline implies, though.

That same hidden reasoning drives the cost story. 68% of billed output tokens are thinking the caller never sees, at 15,164 reasoning tokens against 7,232 answer tokens per task. On Artificial Analysis' full index run, $360 of the roughly $548 total went on reasoning. Cheap tokens, expensive thinking habit. It is the reason AI customer service cost never tracks the rate card.

Four footguns worth knowing before you start

None of these are dealbreakers. Each one costs an afternoon if you meet it cold.

  1. tool_choice only accepts "auto". You cannot force a specific tool, and there is no "required" or "none". Anything else comes back as a 400.
  2. Structured output defaults to off. Under strict: false, Meta's tool-calling reference warns that generated arguments "are not guaranteed to validate against" your schema. Validate before you execute anything.
  3. Claude Code needs three separate changes. A base URL with no /v1, all five model aliases repointed to muse-spark-1.1, then ANTHROPIC_AUTH_TOKEN in place of ANTHROPIC_API_KEY.
  4. MCP is claimed but not documented. The launch post says the model generalises to MCP servers and custom skills. The tool-calling reference documents neither of them. Developer-defined tools are the only extensibility path actually shown.

There is also no server-side cap on runaway custom-tool loops. max_tool_calls limits Meta's built-in tools only, worth knowing if you have ever watched an AI agent loop spin. The delegation half of this is covered in subagent orchestration, and the case for testing it yourself sits in agent evals.

What Meta still has not published

Gaps are review material too, and this list runs longer than it should four weeks after launch.

Not publishedWhy it matters
Max output tokensYou cannot size a request ceiling
Knowledge cutoffYou cannot reason about staleness
Parameter countNo architecture detail of any kind
Regional allowlistEverything known comes from blocked users
Data retention durationPaid prompts skip training, retention is unstated
Uptime or availability SLANothing at all
Minimum cacheable prefixThe $0.15 rate leans on an undocumented variable

Meta also does not appear on the official Terminal-Bench leaderboard, so the self-reported score there has no third-party confirmation. My Muse Spark 1.1 overview covers the methodology dispute around that number. It is still unresolved.

Who should buy it, and who should skip

A quadrant mapping short and long tasks against one-answer and many-step work, marking batch classification and single functions as buy, and multi-hour agent runs and huge-document reasoning as skip
A quadrant mapping short and long tasks against one-answer and many-step work, marking batch classification and single functions as buy, and multi-hour agent runs and huge-document reasoning as skip

Buy it if you run high-volume, short-horizon work: classification, tagging, extraction, routing, single-file code generation, anything where one call produces one answer and you make millions of those calls. The speed is real and the cached rate is excellent, and at roughly $0.29 per task the savings against Opus 5 pricing are big enough to matter.

Skip it if your workload is long-running agents, deep research runs, or anything that has to reason across a large document set. The independent agentic numbers are not close. This is the one case where paying for GPT-5.6 or Opus 5 comes out cheaper once you count the retries. For coding harnesses specifically, Codex is the better-documented path today.

Wait if you are outside the US, or you need a published retention policy, or you need open weights. On that last one, Kimi K3 is the closest near-frontier option with weights you can actually hold. The closed-weights turn is the part of this launch the community has forgiven least.

The moderate read, from someone tracking the leaderboard, is fair enough that I would sign it:

Hacker News

"New respect for Meta Muse Spark. It seems to sit at a lot of sweet spots in the leader board. It's not the best at anything in particular, but it balances cost and performance quite well."

What this changes for a support queue

Almost nothing. I say that as someone whose job is partly to notice when a new model does change something.

Every few weeks a cheaper, faster model ships and somebody asks whether it rewrites the plan for their helpdesk. The honest answer is that the model was never the constraint. In eesel's own cross-validated trials, when agents rewrote an AI draft, roughly 65% of the edits were length and tone. About 20% needed data the AI could not reach in an ERP or logistics system. Only around 5% were the AI being factually wrong. A better model addresses that last 5%. The rest is prompt engineering, retrieval, agent coaching on your team's own sent replies, and integration depth.

Which is also why copilot-style drafting is where most teams should start, the pattern behind agent assist tools, and why a clean escalation path matters more than a leaderboard position does. If you are building the business case rather than the stack, AI vs human cost is the more useful frame.

The failure I watch for hardest in production is the one no benchmark on Meta's launch page measures. An agent narrating a search it never ran. Reporting files it never saved. An agent that claims to have done the work is a harder problem than an agent doing the work badly, and a model that forgets its own reasoning between turns is not the one I would trust to self-report.

Try eesel for support, not a raw model key

If what you actually want is an AI answering customer tickets, the useful question has little to do with which model tops the index this month. It is whether you can prove the thing is safe before it replies to anyone.

That is the part eesel is built around. You can run an AI agent in simulation against your own historical tickets, and read its answers on real past conversations before a single customer sees one. If the answers are not there yet, you start in copilot mode, where it drafts and your team still sends.

Setup is a helpdesk connection, not a project. It plugs into Zendesk in a few minutes, reads the help centre you already wrote, and bills per resolved ticket instead of per million tokens, so the invoice tracks work done rather than how verbose the model felt that day. Free to try, and the pricing is public.

The verdict

Muse Spark 1.1 is a good model wearing the wrong label. Meta sold an agent, and the agent benchmarks are its worst results. On the evidence it is the best fast-and-cheap single-shot model available right now, with a strong code-generation score and a cached rate nobody else matches.

Judge it on the job you have rather than the category Meta filed it under, and it becomes easy to place. Buy it for the millions of short calls. Keep something stronger on the long runs. And do not let a 1M-token window talk you out of retrieval. At roughly an eighth of Opus 5's cost per task, being second-best at the right jobs is a perfectly good business.

Frequently Asked Questions

Is Meta Muse Spark 1.1 any good?
It depends entirely on the job, more than with any other model I have reviewed this year. On static code generation it scores 58.2% on SciCode, second of eight and ahead of Claude Opus 5 at 55.7%. On agentic knowledge work it lands 18th of 19 on AA-Briefcase. Buy it for short, high-volume, one-answer work. For anything long-running, Claude Opus 5 is still the safer pick.
How does Muse Spark 1.1 compare to Claude Opus 5?
Artificial Analysis scores Muse Spark 1.1 at 50.6 on its Intelligence Index against 60.7 for Claude Opus 5, a 10.1-point gap, and Muse Spark does not make the top 20 at all. The trade is cost: $0.292 per index task against $2.336, so roughly eight times cheaper. If you are picking a tier rather than a winner, Opus 5 vs Sonnet 5 walks through the same decision.
Is Muse Spark 1.1 fast?
It is the fastest model on the board at 217.4 output tokens per second, roughly four times Claude Opus 5. The caveat is which clock you read. Its 2.89s time to first token is best in class, but 9.2 seconds of hidden thinking sits behind it, so the first useful token arrives at 12.09s. That still beats Opus 5 comfortably end to end.
Does Muse Spark 1.1 hallucinate?
More than its peers, by two independent measures. Meta's own Petri 3.0 alignment run names input hallucination as the model's primary weakness at 1.67, worse than GPT-5.5 at 1.22 and Opus 4.8 at 1.34. Artificial Analysis separately scores it lowest on the board for knowledge reliability. For a support queue that matters less than grounding, which AI hallucinations in support covers properly.
Can I use Muse Spark 1.1 with Claude Code or Codex?
Yes, and Meta documents all three harnesses, but there are real footguns. Claude Code needs the base URL without /v1, all five model aliases repointed, and ANTHROPIC_AUTH_TOKEN rather than an API key. The bigger trap is that the Chat Completions path drops the model's reasoning between turns, so it repeats and contradicts itself on long runs. Use the Responses API instead.
Is Muse Spark 1.1 worth it for customer support?
As a raw model it is a reasonable, cheap engine for triage and tagging. As a support solution it is not one, because the model was never the hard part. Routing, retrieval, escalation and integration depth are, which is why AI ticket triage and a clean transfer to human move resolution rate more than any model swap.
Where can I use the Meta Model API?
The public preview is US-only, and developers reported being blocked from Canada, Argentina and Vietnam on launch day. Meta has never published the regional allowlist. It reached OpenRouter about a week after launch, which is when hands-on testing outside the US really started. Full rate card and limits are in my Muse Spark 1.1 pricing breakdown.

Share this article

Kurnia Kharisma Agung Samiadjie

Article by

Kurnia Kharisma Agung Samiadjie

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
An AI agent reaching out of a monitor to operate app windows and documents while two colleagues watch, in Meta's blue brand colour
Trending

Meta Muse Spark 1.1: what it is, what it costs, where it loses

Meta's first paid model API ships a 1M-context agent model at $1.25/$4.25. What Muse Spark 1.1 is actually good at, and the benchmarks Meta left off the slide.

Alicia Kirana UtomoAlicia Kirana UtomoAug 5, 2026
Two people reading a large usage meter and adjusting a stack of billing dials, in Meta's blue brand colour
Trending

Meta Muse Spark 1.1 pricing: the bill has four meters

Muse Spark 1.1 lists at $1.25 in and $4.25 out per million tokens. Four separate meters decide your real bill, and the sticker is the smallest of them.

Rama Adi NugrahaRama Adi NugrahaAug 5, 2026
Skywork AI review illustration showing a super-agent turning one prompt into slides, docs and websites
Trending

Skywork AI review (2026): capable agent, messy billing

An honest Skywork AI review: the super-agent makes real slides, docs and websites, but the trial-to-paid billing is where users get burned.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 20, 2026
Editorial illustration of parallel coding agents and terminal windows, representing the field of ZCode alternatives
Trending

The 8 best ZCode alternatives in 2026

ZCode is impressive and rough at the same time. Here are the 8 best ZCode alternatives in 2026, with real pricing, honest trade-offs, and who each one is for.

Rama Adi NugrahaRama Adi NugrahaJul 12, 2026
Editorial illustration of a benchmark leaderboard with one tall highlighted bar, representing ZCode and the GLM-5.2 model
Trending

ZCode: what Z.ai's new AI coding agent really is

A hands-on read on ZCode, the free agentic coding app from the GLM team: the GLM-5.2 model behind it, the real launch-week complaints, and who should use it.

Rama Adi NugrahaRama Adi NugrahaJul 12, 2026
PromptQL alternatives cover banner on an indigo backdrop
Trending

8 best PromptQL alternatives in 2026

The 8 best PromptQL alternatives in 2026, from Databricks Genie to open-source Wren AI, with real pricing, strengths, and who each one is actually for.

Rama Adi NugrahaRama Adi NugrahaJul 10, 2026
PromptQL review cover banner with the PromptQL logo on an indigo backdrop
Trending

PromptQL review (2026): Hasura's data agent, tested

A hands-on PromptQL review: how Hasura's data agent separates planning from execution, what it costs, and whether the reliability pitch holds up.

Alicia Kirana UtomoAlicia Kirana UtomoJul 10, 2026
Shadow, the AI interface for Mac, review cover illustration
Trending

Shadow review (2026): the AI interface for Mac

My hands-on Shadow review: the bot-free AI interface for Mac that transcribes meetings on-device, runs custom Skills from a shortcut, and costs $8 a month.

Alicia Kirana UtomoAlicia Kirana UtomoJul 8, 2026
Hand-drawn illustration with the Grok logomark, a support agent, and benchmark and pricing panels
Trending

Grok 4.5: benchmarks, pricing, and what it means for support

xAI just shipped Grok 4.5. I dug into the real benchmarks, the token pricing, and whether a hot new model actually changes anything for your support queue.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 9, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free