Meta Muse Spark 1.3 review: a fast, cheap coding model with an asterisk

Rama Adi Nugraha
Written by

Rama Adi Nugraha

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 8, 2026

Expert Verified
Illustrated hero banner for a Meta Muse Spark 1.3 review, showing a benchmark scorecard, a coding agent, and a pricing split

What "review" means here, and what I couldn't test

I build integrations and APIs for a living, so when I read a model launch I am less interested in the press release and more in what the thing does once you wire it up. That is the lens for this review. I read Meta's research blog and evaluation methodology closely, pulled the numbers off Artificial Analysis, and read the whole Hacker News launch thread where 261 comments landed in a day.

What I could not do is run max end to end, and that is not a shortcut on my part. Meta's own blog says the previously available reasoning modes ship today "with max reasoning coming shortly after we finish additional safety testing." So the mode that posts the 62 is not the mode most people can call yet. When I talk about live behavior I mean xhigh, the tier that is actually available, which posts a 61. Keep that gap in mind for the rest of the piece, because Meta's marketing does not.

The one chart Meta wants you to read

Every Muse Spark 1.3 review is going to open with this scorecard, so let me put it up and then tell you what it does not say.

Meta's official launch scorecard comparing Muse Spark 1.3, Muse Spark 1.2, GPT-5.6 Sol, and Opus 5, as taken from Meta
Meta's official launch scorecard comparing Muse Spark 1.3, Muse Spark 1.2, GPT-5.6 Sol, and Opus 5, as taken from Meta

The four columns are Muse Spark 1.3, Muse Spark 1.2, GPT-5.6 Sol, and Opus 5, across agent, coding, instruction-following, and long-context buckets. It reads as a clean win for 1.3. Here is the asterisk: the 1.3 column is max, and the 1.2 column it is compared against is xhigh. That is not like for like. Some of the jump you are looking at is a reasoning-tier change, not a model-generation change. An HN commenter, sunaookami, flagged that the scorecard "only shows max reasoning" while max is exactly the gated mode, and called it dishonest. You are being shown the best possible version of 1.3 against a mid version of 1.2.

That does not make the model bad. It makes the chart a marketing artifact, which is what launch charts always are. Read it for direction, not for the exact size of the gap.

Where it genuinely wins: long context and coding

Strip out the framing and the real wins are still real, and they cluster in two places.

On long context, 1.3 is excellent. On Meta's MRCR recall test it posts 98.5 at the 256K-512K band and 98.1 at 512K-1M, well ahead of the 1.2 numbers next to it. If your workload is "read this enormous codebase or document set and reason over all of it," that recall is a legitimate reason to look. The 1M-token context window is not just marketing headroom here; the recall scores say it actually uses the window.

On coding, the deltas versus 1.2 are the most credible part of Meta's blog because they are specific. Meta's engineers report roughly 20% fewer tool calls and 25% fewer tokens on the same coding tasks, with a cleaner, less rambly style. On DeepSWE v1.1 it posts 75.4 against 1.2's 55.0, and it ties GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1. A community reviewer, majerep, called the previous version "the best free model available on OpenCode" for precise, moderate tasks, and the coding numbers say 1.3 extends that lead. If I were reaching for a cheap model to drive an agentic coding loop, this would be on the shortlist next to Kimi K3 and DeepSeek V4 Flash.

Where the "agentic" story is really a coding story

Here is the finding I did not expect walking in. Meta's headline is "improved performance across agentic and coding tasks," and the agentic half of that does not hold up against the model Meta itself chose to compare against.

Where Muse Spark 1.3 leads on the scorecard and where Opus 5 still holds the edge
Where Muse Spark 1.3 leads on the scorecard and where Opus 5 still holds the edge

Walk the six agent rows on Meta's scorecard. Opus 5 max wins GDPVal-AA v2 (1824 to 1754), JobBench (65.7 to 64.9), OSWorld 2.0 (68.3 to 66.9), and AutomationBench (50.3 to 49.4). GPT-5.6 Sol takes DeepSearchQA (93.0) and the Agentic Instruction-Following Index (60.5). That is Muse Spark 1.3 losing every agent row to one rival or the other. Where it dominates is coding and long context, which is a different claim than "best agent." The honest read is that "improved agentic" is mostly "improved coding," and coding is genuinely improved.

I want to be fair here, because this is still a top-ten model and the agent gaps are narrow. But if you are picking a model specifically to drive multi-step tool use and computer control, the scorecard Meta published is quietly telling you Opus 5 is the safer pick for that job. For support automation, where a wrong "agentic" action means a wrong reply to a real customer, that distinction matters a lot, and it is why we simulate every agent against past tickets before it goes live rather than trusting a benchmark row.

The verbosity tax nobody puts on the scorecard

Speed and price look great until you count the tokens. To finish the Artificial Analysis Intelligence Index, Muse Spark 1.3 emitted 120 million output tokens, against a field median of 72 million. That ranks it 139th of 636 on verbosity, meaning it is chattier than most of the field.

Muse Spark 1.3 emitted 120 million output tokens to finish the AA Intelligence Index, against a field median of 72 million, roughly triple Muse Spark 1.2
Muse Spark 1.3 emitted 120 million output tokens to finish the AA Intelligence Index, against a field median of 72 million, roughly triple Muse Spark 1.2

The Hacker News crowd, running their usual pelican-on-a-bicycle SVG stress test, independently pegged it at about 3x the token use of 1.2. Since you pay per output token, verbosity is a direct cost multiplier. A $4.25-per-million output rate that produces three times the tokens is not really a $4.25 model in practice. This is the number I would put next to the pricing table before I signed anything, and it is exactly the kind of hidden total cost that a raw token meter hides and a flat per-resolution price does not.

The pricing is two prices, and the gap is the whole story

Muse Spark 1.3's pricing is the most-discussed thing about it, and rightly so, because it is really two prices with a data toll in between.

Muse Spark 1.3's two pricing endpoints: a private standard endpoint and a much cheaper contributor endpoint that trains on your data
Muse Spark 1.3's two pricing endpoints: a private standard endpoint and a much cheaper contributor endpoint that trains on your data
EndpointInput / 1MOutput / 1MMeta trains on your data?
Standard (xhigh)$1.25$4.25No
Contributor~$0.10~$0.20Yes

The standard rate is moderately priced against a $1.75 input median, and Artificial Analysis measured $0.55 per Index task on it. The contributor endpoint is 10-20x cheaper because you are paying Meta in data instead of dollars. The community read on why the split exists is sharper than Meta's own framing:

Hacker News

"It's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that."

That is the whole thing in two sentences. The premium is not for a better model; it is the price of keeping Meta out of your traffic. For a hobby project that is a shrug. For anything touching customer PII, the "cheap" endpoint is not really an option, so the number you should plan around is the standard $1.25 / $4.25, times the verbosity multiplier above.

What the community actually thinks

The launch-day sentiment was genuinely mixed, and the mix is instructive. The optimists like the price-to-intelligence ratio. As HDBaseT noted, the contributor endpoint is "by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too." The skeptics were not sold on capability, saying of the prior version:

Hacker News

"Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable."

And mromanuk described bouncing off 1.2 on a real web app and going back to Claude, Kimi K3, or DeepSeek V4, while hoping 1.3 clears the bar. The consistent thread: fast and cheap is real, top-tier capability is contested, and the tool you drive it with matters. Several commenters, including meric_, stressed that you should run it in Meta's own Muse Code harness, which it was co-trained on, or you leave efficiency on the table.

Who should run it, and who should skip

Let me be concrete, because a review that ends in "it depends" wasted your time.

  • Run it if your workload is coding or long-context reasoning, you can live in Muse Code, and your data is not sensitive so the contributor endpoint is on the table. At that intersection it is a strong, cheap pick.
  • Consider it for cost-sensitive coding agents where you would otherwise reach for Kimi K3 or DeepSeek V4 Flash. Budget for the verbosity.
  • Skip it, for now if you need best-in-class agentic tool use (the scorecard points at Opus 5), if you need max today (still gated), or if you are handling regulated data and cannot pay the standard premium.
  • Skip the raw model entirely if what you actually want is customer support automation. That is a different product, and I will explain why next.

A model is not a coworker: where this leaves a support team

Here is the trap I watch support leaders fall into every time a model like this trends. They read "top-ten agentic model, dirt cheap" and think they can point it at their inbox. A raw model is infrastructure, like a database or a payments API. It is powerful and it is generic. It does not know your refund policy, it is not connected to your helpdesk, it has no memory of last week's tickets, and it has no built-in brake before it sends a confident wrong answer to a paying customer. Meta even says 1.3 is better at "confirming before consequential actions," which is an admission that a raw model needs that guardrail bolted on.

That gap between a model and a working teammate is the thing eesel exists to fill. eesel is an AI teammate platform, and you hire ready-to-work teammates for specific jobs. The current roster is an AI support agent and an AI blog writer. The support agent shows up already knowing how to do the job: it trains on your knowledge base and past tickets, joins your existing queue, drafts or auto-resolves, and escalates cleanly when it should. The model underneath is a swappable component; a frontier model like Muse Spark 1.3 is exactly the kind of engine it can run on, but the value is everything wrapped around it.

The part that maps directly to a coding-model launch like this one is the eesel CLI. If you like Muse Code because you can drive an agent from a terminal, you will get why this matters. The CLI is an agent-friendly way to operate the same eesel teammate and workspace that lives in the dashboard: a person can run it by hand from a terminal, scripts can automate it in CI, and coding agents such as Claude Code, Codex, and Cursor can drive it directly. It is the same teammate, not a separate cut-down product, so you can manage sources, trigger runs, and pull activity without clicking through a UI. That is the difference between "here is a model key, good luck" and "here is a coworker you can also script."

Try eesel instead of a raw model key

If you got to this review because you are weighing whether to build support automation on a cheap frontier model, the honest answer is that the model is the easy part and the least of your problems. eesel gives you the AI support agent already wired for the job: it runs on frontier models like this one, connects to Zendesk, Freshdesk, and the rest of your stack, and, the part that actually de-risks it, lets you simulate it on thousands of your past tickets so you see its real resolution rate before it ever replies to a customer.

The eesel homepage, showing AI teammates that plug into your existing tools

That simulation step is the lesson we learned the hard way after years of putting AI on live support queues: we have watched confident-sounding bots quietly give wrong answers, which is why nothing at eesel ships to a customer without a dry run against your own history first. You can start free and see the number on your own data.

The verdict

Muse Spark 1.3 is a real step up for coding and long context, and at contributor pricing it is one of the best value-per-token options on the board. As a "top agentic model" it is oversold: Meta's own scorecard shows Opus 5 winning the agent rows, the flagship max tier is gated, and the 3x verbosity means the cheap sticker price is not the price you pay. Read it as what it is, a fast, cheap, verbose coding specialist with a strong long-context habit, and it is easy to like. Just do not confuse a strong model with a working teammate. If the job is answering customers, hire the teammate, not the engine.

Frequently Asked Questions

Is Meta Muse Spark 1.3 good?
For coding and long-context work, yes. It posts an Artificial Analysis Intelligence Index of 62 for the max tier, ranking #6 of 636 models, and it leads Meta's own scorecard on coding and 256K-1M context recall. The catch is that max reasoning was still gated behind safety testing at launch, so the tier you can actually run is xhigh, and it burns roughly 3x the tokens of Muse Spark 1.2. If you want an AI that works your support queue instead of a raw key, an AI agent layered on top of a model is the better fit.
How much does Meta Muse Spark 1.3 cost?
The standard xhigh endpoint is $1.25 per million input tokens and $4.25 per million output tokens, and Meta does not train on that traffic. A separate contributor endpoint runs roughly $0.10 / $0.20 per million, about 10-20x cheaper, in exchange for Meta training on your data. The verbosity means your real bill is higher than the sticker rate suggests, which is why teams often prefer predictable per-resolution pricing over raw tokens.
Is Muse Spark 1.3 better than Opus 5?
It depends on the job. Muse Spark 1.3 wins on long-context recall and several coding benchmarks, but Opus 5 max still beats it on four of six agent rows in Meta's own scorecard. See our Claude Opus 5 review for the other side. For customer support specifically, the model matters less than the resolution rate the surrounding system can hit.
Are Muse Spark 1.3's weights open source?
No. Artificial Analysis lists Muse Spark 1.3 as proprietary. It is a closed, frontier line, distinct from Meta's open Llama family, and no weights release was announced at launch.
Can I use Muse Spark 1.3 for customer support?
You can call it through the Meta Model API, but a raw model does not know your help center, plug into your helpdesk, or hold back on a risky reply. Those are jobs for an AI support agent that you can train on your knowledge base and simulate before it goes live.

Share this article

Rama Adi Nugraha

Article by

Rama Adi Nugraha

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
A developer workbench where one AI model finishes one task cleanly and stalls on the one beside it, in Meta's blue brand colour
Trending

Meta Muse Spark 1.1 review: a high ceiling and a low floor

Muse Spark 1.1 is the fastest model on the board and one of the weakest agentic performers on it. A review of which jobs those cheap tokens actually survive.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 5, 2026
A developer looking at a terminal window with sealed glass workspaces branching off it and a long paper log unspooling underneath
Trending

Meta Muse Code review: the harness is the product, not the model

A hands-on read of Meta's terminal coding agent. The isolation and the audit log are the best parts, and every benchmark gain Meta showed was measured inside Meta's own harness.

Alicia Kirana UtomoAlicia Kirana UtomoAug 18, 2026
A ranked leaderboard column with one card highlighted partway down, and two routes branching away from it toward a cluster of frontier model cards and an open-weights repository box, in Meta's blue brand colour
Trending

Meta Muse Spark 1.2 alternatives: 8 models worth switching to in 2026

Nothing on the Artificial Analysis board beats Muse Spark 1.2 for less money. So the real reason to leave is the weights Meta promised and has not shipped.

Rama Adi NugrahaRama Adi NugrahaAug 18, 2026
A developer at a terminal while a parent agent fans work out to three subagent cards, each with its own branch graph, next to an event log and a benchmark chart, in Meta's blue brand colour
Trending

Meta Muse Spark 1.2: what changed, what it costs, and the catch

Meta shipped Muse Spark 1.2 as a coding release. The coding scores barely moved. The agent scores jumped. Here is what actually changed, and what the cheap tier costs you.

Alicia Kirana UtomoAlicia Kirana UtomoAug 13, 2026
An AI agent reaching out of a monitor to operate app windows and documents while two colleagues watch, in Meta's blue brand colour
Trending

Meta Muse Spark 1.1: what it is, what it costs, where it loses

Meta's first paid model API ships a 1M-context agent model at $1.25/$4.25. What Muse Spark 1.1 is actually good at, and the benchmarks Meta left off the slide.

Alicia Kirana UtomoAlicia Kirana UtomoAug 5, 2026
Illustrated hero banner for a Meta Muse review, showing a personal AI agent working inside a secure cloud VM on travel and shopping while its owner relaxes
Trending

Meta Muse review: is Meta's personal AI agent worth your accounts?

A hands-on Meta Muse review: what Meta's new personal AI agent actually does, the Secure VM and Sentinel security model, the $20 and $100 pricing, and the catch.

Alicia Kirana UtomoAlicia Kirana UtomoSep 9, 2026
Illustration of the IBM Granite 4.2 open model family with reasoning, speech, and security icons
Trending

IBM Granite 4.2 review: is IBM's open reasoning model worth it?

A hands-on IBM Granite 4.2 review: what changed, the benchmarks, real access and pricing, and where the 3B/8B/30B open models fit for support and AI teams.

Alicia Kirana UtomoAlicia Kirana UtomoAug 30, 2026
Editorial illustration of a benchmark leaderboard with one tall highlighted bar, representing ZCode and the GLM-5.2 model
Trending

ZCode: what Z.ai's new AI coding agent really is

A hands-on read on ZCode, the free agentic coding app from the GLM team: the GLM-5.2 model behind it, the real launch-week complaints, and who should use it.

Rama Adi NugrahaRama Adi NugrahaJul 12, 2026
Editorial illustration of parallel coding agents and terminal windows, representing the field of ZCode alternatives
Trending

The 8 best ZCode alternatives in 2026

ZCode is impressive and rough at the same time. Here are the 8 best ZCode alternatives in 2026, with real pricing, honest trade-offs, and who each one is for.

Rama Adi NugrahaRama Adi NugrahaJul 12, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free