
What "review" means here, and what I couldn't test
I build integrations and APIs for a living, so when I read a model launch I am less interested in the press release and more in what the thing does once you wire it up. That is the lens for this review. I read Meta's research blog and evaluation methodology closely, pulled the numbers off Artificial Analysis, and read the whole Hacker News launch thread where 261 comments landed in a day.
What I could not do is run max end to end, and that is not a shortcut on my part. Meta's own blog says the previously available reasoning modes ship today "with max reasoning coming shortly after we finish additional safety testing." So the mode that posts the 62 is not the mode most people can call yet. When I talk about live behavior I mean xhigh, the tier that is actually available, which posts a 61. Keep that gap in mind for the rest of the piece, because Meta's marketing does not.
The one chart Meta wants you to read
Every Muse Spark 1.3 review is going to open with this scorecard, so let me put it up and then tell you what it does not say.

The four columns are Muse Spark 1.3, Muse Spark 1.2, GPT-5.6 Sol, and Opus 5, across agent, coding, instruction-following, and long-context buckets. It reads as a clean win for 1.3. Here is the asterisk: the 1.3 column is max, and the 1.2 column it is compared against is xhigh. That is not like for like. Some of the jump you are looking at is a reasoning-tier change, not a model-generation change. An HN commenter, sunaookami, flagged that the scorecard "only shows max reasoning" while max is exactly the gated mode, and called it dishonest. You are being shown the best possible version of 1.3 against a mid version of 1.2.
That does not make the model bad. It makes the chart a marketing artifact, which is what launch charts always are. Read it for direction, not for the exact size of the gap.
Where it genuinely wins: long context and coding
Strip out the framing and the real wins are still real, and they cluster in two places.
On long context, 1.3 is excellent. On Meta's MRCR recall test it posts 98.5 at the 256K-512K band and 98.1 at 512K-1M, well ahead of the 1.2 numbers next to it. If your workload is "read this enormous codebase or document set and reason over all of it," that recall is a legitimate reason to look. The 1M-token context window is not just marketing headroom here; the recall scores say it actually uses the window.
On coding, the deltas versus 1.2 are the most credible part of Meta's blog because they are specific. Meta's engineers report roughly 20% fewer tool calls and 25% fewer tokens on the same coding tasks, with a cleaner, less rambly style. On DeepSWE v1.1 it posts 75.4 against 1.2's 55.0, and it ties GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1. A community reviewer, majerep, called the previous version "the best free model available on OpenCode" for precise, moderate tasks, and the coding numbers say 1.3 extends that lead. If I were reaching for a cheap model to drive an agentic coding loop, this would be on the shortlist next to Kimi K3 and DeepSeek V4 Flash.
Where the "agentic" story is really a coding story
Here is the finding I did not expect walking in. Meta's headline is "improved performance across agentic and coding tasks," and the agentic half of that does not hold up against the model Meta itself chose to compare against.

Walk the six agent rows on Meta's scorecard. Opus 5 max wins GDPVal-AA v2 (1824 to 1754), JobBench (65.7 to 64.9), OSWorld 2.0 (68.3 to 66.9), and AutomationBench (50.3 to 49.4). GPT-5.6 Sol takes DeepSearchQA (93.0) and the Agentic Instruction-Following Index (60.5). That is Muse Spark 1.3 losing every agent row to one rival or the other. Where it dominates is coding and long context, which is a different claim than "best agent." The honest read is that "improved agentic" is mostly "improved coding," and coding is genuinely improved.
I want to be fair here, because this is still a top-ten model and the agent gaps are narrow. But if you are picking a model specifically to drive multi-step tool use and computer control, the scorecard Meta published is quietly telling you Opus 5 is the safer pick for that job. For support automation, where a wrong "agentic" action means a wrong reply to a real customer, that distinction matters a lot, and it is why we simulate every agent against past tickets before it goes live rather than trusting a benchmark row.
The verbosity tax nobody puts on the scorecard
Speed and price look great until you count the tokens. To finish the Artificial Analysis Intelligence Index, Muse Spark 1.3 emitted 120 million output tokens, against a field median of 72 million. That ranks it 139th of 636 on verbosity, meaning it is chattier than most of the field.

The Hacker News crowd, running their usual pelican-on-a-bicycle SVG stress test, independently pegged it at about 3x the token use of 1.2. Since you pay per output token, verbosity is a direct cost multiplier. A $4.25-per-million output rate that produces three times the tokens is not really a $4.25 model in practice. This is the number I would put next to the pricing table before I signed anything, and it is exactly the kind of hidden total cost that a raw token meter hides and a flat per-resolution price does not.
The pricing is two prices, and the gap is the whole story
Muse Spark 1.3's pricing is the most-discussed thing about it, and rightly so, because it is really two prices with a data toll in between.

| Endpoint | Input / 1M | Output / 1M | Meta trains on your data? |
|---|---|---|---|
| Standard (xhigh) | $1.25 | $4.25 | No |
| Contributor | ~$0.10 | ~$0.20 | Yes |
The standard rate is moderately priced against a $1.75 input median, and Artificial Analysis measured $0.55 per Index task on it. The contributor endpoint is 10-20x cheaper because you are paying Meta in data instead of dollars. The community read on why the split exists is sharper than Meta's own framing:
"It's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that."
That is the whole thing in two sentences. The premium is not for a better model; it is the price of keeping Meta out of your traffic. For a hobby project that is a shrug. For anything touching customer PII, the "cheap" endpoint is not really an option, so the number you should plan around is the standard $1.25 / $4.25, times the verbosity multiplier above.
What the community actually thinks
The launch-day sentiment was genuinely mixed, and the mix is instructive. The optimists like the price-to-intelligence ratio. As HDBaseT noted, the contributor endpoint is "by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too." The skeptics were not sold on capability, saying of the prior version:
"Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable."
And mromanuk described bouncing off 1.2 on a real web app and going back to Claude, Kimi K3, or DeepSeek V4, while hoping 1.3 clears the bar. The consistent thread: fast and cheap is real, top-tier capability is contested, and the tool you drive it with matters. Several commenters, including meric_, stressed that you should run it in Meta's own Muse Code harness, which it was co-trained on, or you leave efficiency on the table.
Who should run it, and who should skip
Let me be concrete, because a review that ends in "it depends" wasted your time.
- Run it if your workload is coding or long-context reasoning, you can live in Muse Code, and your data is not sensitive so the contributor endpoint is on the table. At that intersection it is a strong, cheap pick.
- Consider it for cost-sensitive coding agents where you would otherwise reach for Kimi K3 or DeepSeek V4 Flash. Budget for the verbosity.
- Skip it, for now if you need best-in-class agentic tool use (the scorecard points at Opus 5), if you need
maxtoday (still gated), or if you are handling regulated data and cannot pay the standard premium. - Skip the raw model entirely if what you actually want is customer support automation. That is a different product, and I will explain why next.
A model is not a coworker: where this leaves a support team
Here is the trap I watch support leaders fall into every time a model like this trends. They read "top-ten agentic model, dirt cheap" and think they can point it at their inbox. A raw model is infrastructure, like a database or a payments API. It is powerful and it is generic. It does not know your refund policy, it is not connected to your helpdesk, it has no memory of last week's tickets, and it has no built-in brake before it sends a confident wrong answer to a paying customer. Meta even says 1.3 is better at "confirming before consequential actions," which is an admission that a raw model needs that guardrail bolted on.
That gap between a model and a working teammate is the thing eesel exists to fill. eesel is an AI teammate platform, and you hire ready-to-work teammates for specific jobs. The current roster is an AI support agent and an AI blog writer. The support agent shows up already knowing how to do the job: it trains on your knowledge base and past tickets, joins your existing queue, drafts or auto-resolves, and escalates cleanly when it should. The model underneath is a swappable component; a frontier model like Muse Spark 1.3 is exactly the kind of engine it can run on, but the value is everything wrapped around it.
The part that maps directly to a coding-model launch like this one is the eesel CLI. If you like Muse Code because you can drive an agent from a terminal, you will get why this matters. The CLI is an agent-friendly way to operate the same eesel teammate and workspace that lives in the dashboard: a person can run it by hand from a terminal, scripts can automate it in CI, and coding agents such as Claude Code, Codex, and Cursor can drive it directly. It is the same teammate, not a separate cut-down product, so you can manage sources, trigger runs, and pull activity without clicking through a UI. That is the difference between "here is a model key, good luck" and "here is a coworker you can also script."
Try eesel instead of a raw model key
If you got to this review because you are weighing whether to build support automation on a cheap frontier model, the honest answer is that the model is the easy part and the least of your problems. eesel gives you the AI support agent already wired for the job: it runs on frontier models like this one, connects to Zendesk, Freshdesk, and the rest of your stack, and, the part that actually de-risks it, lets you simulate it on thousands of your past tickets so you see its real resolution rate before it ever replies to a customer.
That simulation step is the lesson we learned the hard way after years of putting AI on live support queues: we have watched confident-sounding bots quietly give wrong answers, which is why nothing at eesel ships to a customer without a dry run against your own history first. You can start free and see the number on your own data.
The verdict
Muse Spark 1.3 is a real step up for coding and long context, and at contributor pricing it is one of the best value-per-token options on the board. As a "top agentic model" it is oversold: Meta's own scorecard shows Opus 5 winning the agent rows, the flagship max tier is gated, and the 3x verbosity means the cheap sticker price is not the price you pay. Read it as what it is, a fast, cheap, verbose coding specialist with a strong long-context habit, and it is easy to like. Just do not confuse a strong model with a working teammate. If the job is answering customers, hire the teammate, not the engine.
Frequently Asked Questions
Is Meta Muse Spark 1.3 good?
How much does Meta Muse Spark 1.3 cost?
Is Muse Spark 1.3 better than Opus 5?
Are Muse Spark 1.3's weights open source?
Can I use Muse Spark 1.3 for customer support?

Article by
Rama Adi Nugraha
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








