
Why people are looking past Muse Spark 1.2 at all
I build integrations and API plumbing at eesel, so with any model launch my instinct is to skip past the launch post and go read the rate card and the rate limits, then the model page. Muse Spark 1.2 is interesting for a precise reason: the marketing and the data disagree here in the opposite direction from the usual one. The launch framing undersold where it landed on price-performance, and it oversold what you'd get to own.
There are three things actually pushing people to look elsewhere.
The weights are a promise, not a download. Meta's Chief AI Officer said that an open-weight version was coming "soon". Meta's verified Hugging Face org reports four repositories at the moment, and all four of them are Muse Glimmer. Search Muse-Spark across Hugging Face and you get back zero results, from anyone.
Nobody can independently check the model. For most of the board, Artificial Analysis publishes output speed and latency. For Muse Spark 1.2 those two fields are just blank. Muse Glimmer sits on the same table and carries them. You can price this model, but you cannot benchmark how fast it answers.
The cheap tier costs you your traffic. The muse-spark-1.2-contributor ID is 12.5x cheaper on input, and Meta's own docs describe this as discounted pricing, given in exchange for permission to train on your prompts and your completions. The arithmetic on that trade I already covered in the Muse Spark 1.2 pricing breakdown.

Four days ahead of Meta's announcement, one Hacker News commenter made the case for the weights better than Meta's own post managed to:
"Meta, please go back to publishing weights. Your models aren't top-tier, this wouldn't hurt you at all at your current position in the rankings. They wouldn't have the best open-weight models, but the best western open-weight models."
Meta did exactly this four days later, but for a 30B student model. The frontier model got an adjective instead.
How I picked these 8
I ranked on three things, and in this order.
Independent score, not vendor charts. Every index number on this page is the Artificial Analysis Intelligence Index v4.1.1. It runs nine evaluations, so no lab gets to pick its own comparison set. Which matters here, because the Muse Spark 1.2 launch charts swapped the agent harness between versions.
Cost per task, not sticker price. Verbosity is what per-million-token rates hide. Muse Spark 1.2 burns 30,430 output tokens on each index task, so the real cost sits well above whatever the $1.25 headline implies to you. Cost per task is the number that survives contact with an actual bill.
Whether the weights exist. For every model on the list I went and checked the licence field on the actual repository, rather than trusting a press release. Two of the eight turned out to be licensed differently from the way they usually get described.
One obvious name I left out. There is no eesel model in this list, for the reason that eesel is not a model vendor, and dropping my own employer into a benchmark table would be exactly the kind of thing which makes these posts useless.
The comparison table
Index scores and cost per task both come from the Artificial Analysis models board, checked 18 August 2026. Token rates are from each vendor's own pricing page, checked on the same day.
| Model | Index | Cost / task | Input /1M | Output /1M | Context | Weights | Licence | Best for |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 63.05 | $2.3369 | $5.00 | $25.00 | 1M | No | Proprietary | Highest ceiling on the board |
| GPT-5.6 Sol | 60.93 | $1.2312 | $5.00 | $30.00 | 1.05M | No | Proprietary | Frontier work on OpenAI tooling |
| Grok 4.6 | 60.92 | $0.8367 | varies | varies | 2M | No | Proprietary | Near-frontier at mid cost |
| Kimi K3 | 59.70 | $0.8375 | $3.00 | $15.00 | 1.05M | Yes | Kimi K3 License | Open weights that actually beat Muse Spark |
| Qwen3.8 Max | 58.08 | $1.1320 | $2.00 | $6.00 | 1M | Base only | qwen3.8-max | Hosted power with an open fallback |
| Muse Spark 1.2 | 56.76 | $0.3992 | $1.25 | $4.25 | 1.05M | No | Proprietary | The incumbent you are reading about |
| GPT-5.6 Terra | 56.58 | $0.5080 | $2.00 | $12.00 | 1.05M | No | Proprietary | Like-for-like swap with better tooling |
| Gemini 3.7 Flash | 56.03 | $0.4022 | $0.75 | $3.75 | 1.05M | No | Proprietary | Speed at near-identical cost |
| DeepSeek V4 Pro | 53.20 | $0.2521 | $0.66 | $1.98 | 1M | Yes | MIT | The cheapest permissive licence here |
| GPT-5.6 Luna | 52.32 | $0.0471 | $0.20 | $1.20 | 1.05M | No | Proprietary | The price floor for hosted models |
| Muse Glimmer 30B | 35.06 | $0.0732 | local | local | 131K | Yes | Apache 2.0 | Running a Meta model on your own GPU |
DeepSeek rates shown here are the off-peak ones. Gemini rates shown are introductory, running through 31 December 2026.
Plot it out, cost against capability, and the shape of the decision gets a lot easier to see.

Pick your escape route
Instead of reading through all eight, just answer the one question which actually splits the field.
What do you actually need that Muse Spark 1.2 is not giving you?
Pick one. Every answer costs you something, and the number is shown.
Top of the Intelligence Index at 63.05, which is 6.29 points clear of Muse Spark 1.2. It carries the full 1M context at standard pricing with no long-context step up.
The trade: $2.3369 per task against $0.3992. You are paying 5.9x for roughly an 11% capability gain.
59.70 on the index, and the only open-weights model on the 20-row board that beats Muse Spark 1.2. Weights shipped on the day Moonshot said they would.
The trade: $0.8375 per task, so 2.1x the cost for 2.94 index points. The licence is bespoke, not MIT.
MIT licensed with no bespoke terms, 53.20 on the index, and the cheapest per task of any credible option here at $0.2521. 662 community repositories exist, 316 of them quantized.
The trade: 3.56 index points below Muse Spark 1.2, and the rate doubles during peak UTC hours.
Meta's own Apache 2.0 model, distilled from Muse Spark, running under 20GB at roughly 4-bit. This is the Meta model you can actually download today.
The trade: 35.06 on the index, so 21.70 points below Muse Spark 1.2. A capable local agent, not a frontier substitute.
1. Claude Opus 5
Best for: teams where a wrong answer costs more than the API bill does.
Anthropic's flagship holds the top of the board at 63.05. It is also the only model here where the gap over Muse Spark 1.2 is wide enough that you would notice it without a benchmark harness at all.
Where it beats Muse Spark 1.2. Six index points, and in practice that shows up on the long multi-step tasks, the ones where an early mistake compounds. It carries the full 1M context at standard pricing too, with Anthropic stating plainly that a 900k-token request bills at the same per-token rate as a 9k one does. Neither OpenAI nor xAI is doing that.
Where it doesn't. Cost, and by a wide margin. Then the thing that most coverage skips over: Opus 5's AA-Omniscience hallucination rate rose 14 points to 50%. A higher index score is not the same thing as a more truthful model.
Pricing. $5.00 input, $0.50 cache read, $25.00 output per million. Batch comes at half price. Cache writes are $6.25 at the 5-minute TTL. If that ceiling is more than what you need, Claude Sonnet 5 is now permanently $2.00 and $10.00, this after Anthropic cancelled the increase which had been scheduled for 1 September.
Verdict: the right pick for when a mistake is expensive and the token bill is not. Skip over it if you are running high-volume classification, because there you would be paying 5.9x per task for capability which the task never uses.
2. Kimi K3
Best for: anyone whose actual requirement is "open weights, and better than what I have."
Moonshot's flagship is the single most interesting entry on this list, and the reason is that it is the only model satisfying both halves of what people say they want out of Meta. It scores 59.70, so above Muse Spark 1.2, and the weights are downloadable right now, today.
Where it beats Muse Spark 1.2. Nearly three index points, and then 2.8 trillion parameters that you can host yourself. The part worth sitting with is this: Moonshot's tech blog said the weights would be released by 27 July 2026, and the first public commit on Hugging Face lands 27 July 2026. A named date, met to the day.
Where it doesn't. Two things here. Cost per task is $0.8375, so more than double. Then the licence, which is not what most write-ups claim: the repository's own metadata reads license: other with license_name: "kimi-k3", a bespoke Moonshot licence. Open weights, yes. Permissive open source, no. Its Openness Index is 38.89, which sits below Muse Glimmer's 44.44.
Pricing. $3.00 input, $0.30 cache hit, $15.00 output per million, so Claude Sonnet territory rather than any budget territory. My full Kimi K3 review goes into the reasoning-effort behaviour, which you cannot turn off at any level.
Verdict: the default answer for anybody who was waiting around on Meta's weights. You pay roughly double per task, and you read the licence carefully before you ship, but the thing exists today and it is measurably better.
3. Qwen3.8 Max
Best for: teams that want a hosted model now and a self-hosted fallback later.
Alibaba's flagship scores 58.08 hosted, and this is the entry where my own research went and corrected me mid-draft. Earlier this month the Qwen 3.8 weights were promised and still unpublished, sitting in exactly the position Muse Spark 1.2 is in now. Then they shipped on 9 August.
Where it beats Muse Spark 1.2. 1.32 index points hosted, while the open base Qwen3.8-2.4T-A95B still scores 57.70 on its own. Which means even the downloadable version beats the model you are considering leaving.
Where it doesn't. Cost per task is $1.1320, so nearly 3x. The open release is also not feature-equivalent to the hosted product: Qwen's own model card states that the hosted Max adds vision input, non-thinking support, 1M context by default, plus built-in tools. And the licence is bespoke once again, qwen3.8-max, not Apache.
Pricing. $2.00 input, $0.25 implicit cache read, $6.00 output per million. The explicit cache read is cheaper at $0.17, but it costs $2.50 per million to create, so it only pays off on prefixes you repeat heavily. Practical access routes I put in how to access Qwen3.8 Max.
Verdict: the best hedge on this list, since you can start hosted and then move onto your own hardware without changing model families. The Qwen3.8 Max review carries the human-preference numbers, and those point in a friendlier direction than the automated index does.
4. GPT-5.6 Terra
Best for: a like-for-like swap where the surrounding tooling matters more than the score.
OpenAI's mid-tier lands at 56.58, and that is 0.18 points below Muse Spark 1.2. For practical purposes the two models are the same capability at a different price, which is the reason Meta's launch post compared against Terra rather than GPT-5.6 Sol. One commenter caught this at the time:
"They chose to compare against Open AI's mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games?"
Where it beats Muse Spark 1.2. Ecosystem, mostly. Terra runs inside the tooling your team is probably already using, and OpenAI publishes speed and latency figures which Meta does not.
Where it doesn't. It is 27% more expensive per task, and for a score that is marginally lower. There is a long-context cliff on it too: prompts going over 272K input tokens get priced at 2x input and 1.5x output across the entire request, not only the overage, per OpenAI's model page. Muse Spark 1.2 carries no long-context premium at all.
Pricing. $2.00 input, $0.20 cached, $12.00 output per million, and that is before the long-context multiplier. Where exactly that cliff bites is covered in the GPT-5.6 pricing breakdown.
Verdict: switch for the tooling and for the published latency numbers, but not for the capability. If your prompts routinely run past 272K tokens, do the arithmetic first, before you migrate.
5. Gemini 3.7 Flash
Best for: high-volume workloads where throughput is the constraint.
Google's Flash line scores 56.03 at $0.4022 a task, putting it inside a rounding error of Muse Spark 1.2 on both axes. The difference is that Google publishes how fast the thing is. Artificial Analysis clocks it at 340.1 tokens per second, ranked first out of 188 models measured.
Where it beats Muse Spark 1.2. Measured speed, plus a cheaper sticker rate at $0.75 input. It is the only model here where you can also dial the capability down on purpose: running at low effort scores 50.9 at $0.16 a task, and for classification work that is often the right call.
Where it doesn't. There are two real caveats. The current rate is introductory through 31 December 2026 and it doubles to $1.50 and $7.50 on 1 January 2027, so whatever cost model you build today has a cliff sitting inside it. Its hallucination rate also regressed to 64.5% from the previous version's 55.6%, which I dug into over in the Gemini 3.7 Flash review.
Pricing. $0.75 input, $0.075 cached, $3.75 output per million until the end of 2026. Max output sits at 65,536 tokens, which is notably tighter than the rest of this list.
Verdict: the best pick if latency is the actual complaint you have about Muse Spark 1.2, given that you cannot even measure Meta's. Budget for the January price change now, rather than getting surprised by it later.
6. DeepSeek V4 Pro and V4 Flash
Best for: the cheapest weights you can do what you like with.
This is the entry which wins on licence rather than on score. V4 Pro scores 53.20 at $0.2521 a task, and both it and DeepSeek V4 Flash carry a plain mit licence inside their repository metadata. No bespoke terms. No usage policy sitting alongside it either.
Where it beats Muse Spark 1.2. Price, and freedom. It is 37% cheaper per task, it is MIT licensed, and it has the deepest community ecosystem in this whole comparison: a search returns 662 repositories, with 316 of them quantized. That ecosystem is the thing which makes self-hosting realistic instead of theoretical.
Where it doesn't. It gives up 3.56 index points. There is a quirk worth planning around as well: DeepSeek charges double during peak hours, which are defined as 01:00 to 04:00 and 06:00 to 10:00 UTC. A European support queue is going to hit that window daily; a US one mostly will not. Then on data policy, DeepSeek's paid API terms are silent about training use rather than permissive, and the infrastructure itself sits under PRC law.
Pricing. V4 Pro is $0.66 input and $1.98 output off-peak, and that doubles at peak. V4 Flash comes far cheaper at $0.22 and $0.66 off-peak, scoring 51.77. The peak-hour arithmetic is walked through properly in the DeepSeek V4 Flash pricing post.
Verdict: if the requirement you have is "MIT weights, cheap", then this is the answer and nothing else on the list is competing. If the workload touches regulated customer data, residency is the question to resolve first.
7. GPT-5.6 Luna
Best for: the volume work you should never have been running on a frontier model.
Luna scores 52.32 at $0.0471 a task. That is 8.5x cheaper than Muse Spark 1.2 for 92% of the index score, and it is the one number here that should make anybody running bulk classification reconsider their whole setup.
Where it beats Muse Spark 1.2. Cost, and dramatically. Most production workloads are not frontier-shaped at all: tagging and routing, extraction, summarising, these rarely need 56 points of index. Luna is fully hosted with published latency too, which Meta's model is not.
Where it doesn't. It is 4.44 points down, and that gap does show on anything properly multi-step. A workhorse, this one, not a reasoner.
Pricing. $0.20 input, $0.02 cached, $1.20 output per million, and with the same >272K long-context multiplier that the rest of the GPT-5.6 family carries.
Verdict: the most underrated line on that table. Before migrating everything onto a pricier model, split your traffic up and check how much of it actually needs the expensive one.
8. Muse Glimmer 30B
Best for: running a Meta model on hardware you own, today.
If you came here because Meta weights were what you wanted, this is the thing which actually exists. Glimmer is 29.78 billion parameters under Apache 2.0, distilled out of Muse Spark, with a 131,072-token context and a January 2026 knowledge cutoff.
Where it beats Muse Spark 1.2. You can download it. Meta quantizes down to roughly 4-bit, which puts the language model under 20GB and fits it onto a 24GB card. Cost per task is $0.0732, and this is the only Meta model carrying an Openness Index at all, at 44.44. The community response came immediately: 296 Glimmer repositories now exist, and unsloth's GGUF build at 755,125 downloads outdraws Meta's own weights repository.
Where it doesn't. 35.06 against 56.76 is a 21.70-point gap, and no quantity of local convenience is closing that. There is a licence footnote worth knowing about as well: alongside the real Apache 2.0 text, the repository ships a Meta-authored USAGE_POLICY.md which restricts use by minors and prohibits military and weapons applications, also critical-infrastructure ones. Apache 2.0 imposes none of those terms, so the practical surface here is Apache plus a policy document.
Pricing. Free to download, then your own compute after that. One developer running it on a 32GB Mac mini was candid about what this means in practice:
"I am getting good results with muse-glimmer running locally, with the caveat that everything runs slowly (e.g., give it a task and then go walk outside or do Qi Gong exercises for a while)."
Verdict: a very good local agent model, and a bad Muse Spark substitute. Treat it like a different tool, rather than a downgrade path.

What the community actually says
Sentiment on Muse Spark 1.2 splits cleanly along a single line. People who read the benchmarks attack its ranking; people who actually ran it praise how few tokens the thing wastes. Both sides are right, and that is why the launch reaction reads so contradictory.
The sharpest sceptical read of it arrived on launch day:
"Somewhat surprised that Meta with all their resources couldn't make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. The use traces must be crucial to functionality which is why they're keeping prices so low."
Two weeks of independent measurement has partly answered this. On the current Artificial Analysis board, Muse Spark 1.2 is not dominated on price-performance at all; it sits on the frontier at its own cost point. The second half of that comment, the part about pricing as a data-acquisition play, holds up much better.
On the weights, the objection is not really about this announcement. It is about the previous one:
"Meta in this case, started nicely with Llama, then switched to closed models, kicked out researchers to build data labeler CEO empire inside Meta. Now opening again, what's next? closing again?"
And the position I saw most often was not "switch" and not "stay" either. It was to take the open artefact and then decline the hosted product:
"I'd also argue this is the case for any company releasing open weights. They're not righteous, they're marketing. That's not necessarily a bad thing! [...] I still will not use a hosted Meta product, but damn this model looks solid."
For balance: the loudest positive voice on the announcement was Box's CEO, and he read it as a category shift rather than a product launch:
"Meta releasing Muse Spark 1.2 as open weights is a very big deal. America now finally has its response to the open weights AI race."
That reaction was to the promise, though. Eight days on and the download still does not exist.
What none of these swaps will fix
Here is the part I care about the most, and that is because I spend my working life on the layer sitting underneath these models.
If you are choosing between the eight for a general coding or research workload, the table above is most of what you need. But a large share of the people searching for Muse Spark 1.2 alternatives are trying to fix a support automation which is giving wrong answers, and a model swap is almost never the fix for that.
The example I opened with is the cleanest one that I know of. A B2B technical support team had their AI confidently telling customers "yes, we support your car model" for vehicle brands which were not in their database. The model here was not broken. Their help centre said "we support all models", the AI retrieved that one sentence, and then it answered correctly from the text it had been given. Moving from a 56.8 model up to a 63.1 model changes nothing about that outcome, because the failure happened in retrieval and not in reasoning.
You can watch the same thing happening in the benchmark data. Muse Spark 1.2's AA-Omniscience accuracy actually fell from 1.1, from 52.0% to 45.4%, while the hallucination rate improved from 50.0% to 33.3%. It did not get more knowledgeable. What it got better at is declining. For a general assistant that is a reasonable trade, and for a support agent it is a poor one, because there a declined answer is still a ticket.
There are three things which move support accuracy more than the model does:
- What the AI can see. Past resolved tickets, and not only the help-centre articles. Most tools train on published documentation alone, which is exactly the reason they repeat "we support all models" back at you.
- What it does when it is unsure. Confidence-based routing into a draft rather than a live reply, so a low-confidence answer turns into an escalation instead of a customer-facing mistake.
- Whether you tested it before your customers did. Simulating against your own historical tickets is what tells you the real resolution rate. A public index score only tells you somebody else's.
I go further into all of this in AI models for support, and into the failure pattern itself in AI hallucinations in support.
Try eesel for support work
If the reason you landed on a model comparison is that your support AI keeps answering wrong, then the model tier is the wrong dial to turn. eesel is an AI teammate which plugs into Zendesk, Freshdesk, Gorgias, HubSpot or Front, and it learns from your resolved ticket history on day one, not only from your public knowledge base.
The part that matters for everything above: before it ever touches a customer, you run it in simulation across your own past tickets and see the real resolution rate, broken down by theme. Gridwise saw 73% of tier-1 requests resolved inside the first month. Billing is per resolved ticket at $0.40, and not per million tokens that you cannot forecast, so the peak-hour and the long-context arithmetic on this page stops being your problem.

Free to try, no card needed, and inside a day you will know whether the model was ever your problem at all. Try eesel.
So which one should you actually pick
If you take just one thing away from this page, make it this one: Muse Spark 1.2 is not a bad deal on capability per dollar, and switching for score alone will cost you more than it gives you.
Leave for a concrete reason instead. If a wrong answer is expensive, Claude Opus 5 buys you the biggest real gap. If what you need is weights that beat what you have today, Kimi K3 is the only model doing both, at roughly double the per-task cost. If MIT specifically is the requirement, DeepSeek is the answer and the licence there is clean. And if speed is the complaint, Gemini 3.7 Flash is the only one of the pair whose speed anybody has published.
And if you are still waiting on Meta's weights, the ledger is worth remembering here. Moonshot named a date and then hit it. Alibaba shipped inside a week. DeepSeek has been MIT the whole time. Meta shipped the student model and an adjective.
Frequently Asked Questions
What are the best Meta Muse Spark 1.2 alternatives?
Is Muse Spark 1.2 open source?
muse-spark-1.2 is a proprietary API model with no published weights. Meta announced on 10 August that an open-weight version was coming soon, without naming a licence or a date, and nothing has shipped since. The Meta model that is open is Muse Glimmer 30B, under Apache 2.0. For the wider picture see my roundup of open source AI agents.How does Muse Spark 1.2 pricing compare to the alternatives?
muse-spark-1.2 ID is $1.25 per million input tokens and $4.25 output, which works out to $0.3992 per Intelligence Index task. That is cheaper per task than Kimi K3 pricing at $0.8375 and cheaper than Claude Opus 5 at $2.3369. The full breakdown of Muse Spark 1.2 pricing lives in my pricing teardown.Which open weights model is closest to Muse Spark 1.2?
Should I switch off Muse Spark 1.2 for a customer support use case?
What is Muse Glimmer 30B and can it replace Muse Spark 1.2?
Does Muse Spark 1.2 have a free tier?
muse-spark-1.2-contributor ID costs $0.10 input and $0.20 output in exchange for Meta using your traffic to train future models, and it is capped at 100 requests per minute. If a low bill is the goal without the data trade, GPT-5.6 Luna at $0.0471 per task is the cheapest hosted option on the board.
Article by
Rama Adi Nugraha
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








