
How I graded this review
Seven dimensions, each graded off a primary source rather than a press summary: the licence file, the pricing page, the shipped config.json, the API reference, the tech blog's own benchmark table, the platform FAQ, and the deployment recipes. Where I use a hands-on judgement, it comes from a named person in a public thread, not from me guessing.

For the full teardown of what ships in the box, I wrote a companion piece on LongCat 2.0 itself. This review is the buying decision on top of it.
Can you actually get to it? Pick your route
The unusual thing about reviewing this model is that "is it good" and "can you use it" have different answers depending on how you reach it. Three routes exist and each has a different wall. Pick yours:
That third column is why the MIT licence, which is the model's best feature on paper, does almost nothing for most readers. Plain MIT on a checkpoint you cannot fit is a licence to admire.
What it is actually good at
The single most useful datapoint in this whole review is not a benchmark. It is one developer's account of running 3.6 billion tokens through the model during the two months it was on OpenRouter as a stealth model called owl-alpha:
I used this for over 3.6 billion tokens when it was owl-alpha on Openrouter (with Hermes Agent). It was a very good experience.
It's not as 'smart' as other frontier models when it comes to benchmark style tests (one shots, riddles, etc) but it was very good at (1) following instructions, (2) making a plan, (3) following that plan, and (4) staying coherent at very high contexts. I built a number of apps from start to finish and it performed very well.
That is a precise description of a good agentic model and a mediocre chat model, and it matches the second-best hands-on account, from someone who bought a token pack after launch:
As an aside, I also nabbed a 50m token pack for LongCat 2.0 to give it a whirl. Not free, but it's so cheap they're basically giving it away. Very impressed too [...] Not frontier-level intelligence, but a dependable workhorse that can navigate a codebase well and can reliably execute what you tell it to do.
Two independent users, two months apart, on different access routes, landing on the same verdict: reliable executor, not a genius. That is a more valuable thing to know than any SWE-bench delta, and it is the profile you want if you are building custom coding agents where the harness does the thinking and the model does the work.

Worth flagging the caveat the same user raised, because it changes how you read the benchmark table: LongCat 2.0 is not a reasoning model. Its scores were set without an extended thinking budget, while several models it is compared against had one. That cuts both ways, and it is the kind of asymmetry that makes cross-vendor comparisons in agentic coding CLIs harder than the marketing charts imply.
What it is not good at
The dissent is real and I am not going to bury it. From someone using it in production:
I was using Owl Alpha a lot for my project. Thats not gpt 5.5 level model. it not even close to flash 2.5 model - its not gollowing promts.
And a blunter one, on code specifically:
I have been trying to use longcat 2 but its bad model, it cant follow orders for example. Its coding is terrible, buggy as hell, stay away. Deepseek is way better.
Both are low-vote comments and both contradict the high-token accounts, so I weight them accordingly. But the pattern across all of them is consistent: this model's quality depends heavily on the harness driving it. The people reporting success were running it inside a real agent loop. The people reporting failure were mostly prompting it directly.
On the reasoning side, the most-discussed public test on Hacker News put it third:
Overall I rate Gemini Flash the best, Qwen 3.7 Plus an acceptable second, and LongCat-2.0. an ok'ish third, if you have nothing better.
I would cite that carefully, because the test itself got picked apart in-thread by three separate commenters who argued the question was leading or had no single right answer. It is one prompt, not an eval.
Two integration failures reported on launch day are more actionable than any of the above. A user could not get tool calls working at all because the model emits a <longcat_tool_call> wrapper their harness did not recognise, and another asked a question in English with search enabled and got answers back in Chinese. Neither is a quality problem. Both are the kind of thing that eats an afternoon.
The benchmarks, read honestly
Meituan publishes SWE-bench Pro, not Verified, and it ran the numbers in-house on Claude Code with a 4c8g sandbox at temperature 1.0, stating that "problematic tasks corrected." Here is the field, with every score's origin noted, because the deltas under a few points are noise:
| Model | SWE-bench Pro | Who ran it | Note |
|---|---|---|---|
| Claude Fable 5 | 80.0 | reported by OpenAI | not on an Anthropic page |
| Claude Opus 4.8 | 69.2 | reported by four vendors | the most consistent figure in the set |
| Qwen3.8-Max | 67.7 | Qwen in-house | also states it corrected the task set |
| GPT-5.6 Sol | 64.6 | OpenAI in-house | n/a |
| GPT-5.6 Terra | 63.4 | OpenAI in-house | n/a |
| GPT-5.6 Luna | 62.7 | OpenAI in-house | n/a |
| GLM-5.2 | 62.1 | Z.ai, OpenHands, 400K ctx | tailored prompt |
| LongCat 2.0 | 59.5 | Meituan in-house, Claude Code | no reasoning mode; task set corrected |
| GPT-5.5 | 59.4 | OpenAI in-house | LongCat's own card lists this as 58.6 |
| MiniMax-M3 | 59.0 | MiniMax in-house | n/a |
| DeepSeek-V4-Pro (Max) | 55.4 | DeepSeek in-house | architectural near-twin of LongCat |
| Gemini 3.1 Pro Preview | 54.2 | reported by OpenAI | n/a |
| DeepSeek-V4-Flash (Max) | 52.6 | DeepSeek in-house | 284B model |
Read plainly: 59.5 puts LongCat 2.0 level with last year's GPT-5.5, a few points behind GLM-5.2 and the whole GPT-5.6 family, and roughly ten points behind Opus 4.8. It beats its closest architectural twin, DeepSeek V4 Pro, by four points, which is the comparison I would actually make.
Two things the table does not show. Meituan's own comparison set contains no open-weight rival at all, which r/LocalLLaMA noticed immediately:
I don't know why they won't line up their benchmark to other Chinese and open models, you know, they have DeepSeekV4Pro, KimiK2.7-Coder, GLM5.2, MiniMaxM3, Qwen3.5-397B, MiMoV2.5-Pro.
That is a fair hit. A vendor that benchmarks only against models it beats is telling you where it chose to stand. And Meituan publishes no SWE-bench Verified score at all, so any comparison against the 79 to 81 band that DeepSeek V4 and MiniMax report on Verified is not a comparison you can make.
Price is the strongest column, with a footnote
| List | Promo | Change | |
|---|---|---|---|
| Input, uncached | $0.75 / 1M | $0.30 / 1M | -60% |
| Cache read | $0.015 / 1M | $0.006 / 1M | -60% |
| Output | $2.95 / 1M | $1.20 / 1M | -59% |
Two details from the pricing page that matter more than the headline. First, there is no context-length tiering at any length, which is unusual: Gemini, GPT-5.6 and MiniMax all step the price up past a threshold, and a 200K-token request here costs the same per token as a 2K one. Second, the cache-write price is absent from the page, which is not the same as free.
The footnote is the promo itself. It is labelled limited-time with no published end date. At list price the model is more expensive than DeepSeek V4 Pro's $0.435 and $0.87 for what is, on paper, the same shape of model: 1.6T total, roughly 48 versus 49 billion active, MIT licence. So the entire cost argument rests on a discount the vendor can end whenever it likes. Someone on Reddit had already set the threshold before launch, and the promo cleared it:
as long as it stays below .40 input and .80 output it will have a use.
Note that at list price, it does not clear it. And the "cheapest 1M-context model" line that circulated at launch was corrected in-thread almost immediately, because DeepSeek V4 Flash is $0.14 and $0.28. If your only criterion is price per token, LongCat 2.0 is not the winner even inside the Chinese open-weight field. For the wider field, the GPT-5.6 pricing and Claude Opus 5 pricing pages set the ceiling, and Kimi K3 pricing at $3 and $15 shows that open weights and cheap tokens do not always travel together.
Three walls the marketing does not mention
The context window is 256K, not 1M
Every launch write-up says 1M. The shipped config.json on Hugging Face caps max_position_embeddings at 262,144, with YaRN provisioned up to 983,040 behind that cap. The hosted API adds a separate 131,072-token output ceiling, and max_tokens counts against your context. So the honest number is 256K in, 128K out, and the 1M figure describes the training data.
This is a smaller gap than it sounds, since 256K is a lot, but it is the difference between "fits the monorepo" and "does not." If long context is your actual requirement, compare against how Claude Code's context window behaves in practice rather than against a spec sheet.
The tooling contract is not written down
This is the grade that surprised me most. The API reference publishes a handful of request parameters and no tools array in either the OpenAI-format or Anthropic-format body schema. There is no tool_choice, no parallel_tool_calls, no stop_sequences, no metadata, and content is documented as a plain string, so no content blocks, no images, no tool_result. temperature runs 0 to 1 rather than OpenAI's 0 to 2, which will silently clip requests ported from another provider.
The platform's "tools" page is not a tool catalogue at all. It is a compatibility list of twelve third-party coding clients you can point at the API. There is no server-side web search, no code interpreter, no retrieval store, no MCP surface, and no hosted agent loop. Everything agentic is client-side by design. That is a legitimate architecture, and it is the same bring-your-own-harness bet as GPT-5.1-Codex-Max in a different form, but it means the model's agentic AI story is entirely your harness's story.
There is no paper, and barely a repo
meituan-longcat/LongCat-2.0 on GitHub is a README, a licence and three figures, roughly 1 MB, zero releases. The weights are on Hugging Face. There is no technical report, only a blog post, which means the two new architecture pieces have no published methodology. The model card omits model_type, so Transformers AutoModel fails outright, and SGLang is the only engine, with its support PR closed unmerged and a nightly wheel required.
For a model whose whole pitch is openness, that is thin. The comparison that makes it obvious is Hugging Face itself: 3,240 downloads a month and zero inference providers on the hub is not the footprint of a model people are deploying.
The grade that stops enterprise buyers
Here is the finding I would put in front of any security reviewer. Meituan's platform FAQ is silent on data retention, training on prompts, data residency, and SLA. The word "training" does not appear anywhere on it.
That silence is not a technicality. It is the exact question every buyer asks, and I hear it on almost every call. A technical evaluator at a B2B hardware company I spoke with in March would not proceed until they had a straight answer on whether the AI could reach outside their approved knowledge; a separate buyer, gated by an internal security review, needed written assurance that ticket data containing card numbers and passwords stayed inside their environment. Those are not exotic asks. They are the floor. A vendor page that does not mention retention cannot clear that floor, whatever its benchmark score is.
Developers made the same call independently. One skipped the model entirely during its free stealth period on policy grounds, then found out afterwards what he had passed up:
Wait, this is Owl Alpha? Now I wish I had tried it when it was available. I stayed away from it back then because of their privacy policy
And the OpenRouter workaround does not fix it, it just documents the exposure: AtlasCloud is the only provider, it carries no zero-data-retention badge, and its policy states 7-day content retention. If you are working through a SOC 2 and GDPR review, or anywhere near HIPAA-compliant AI requirements, that is where this ends. Self-hosting is the only route that removes the question, and self-hosting means eight B300s.
To be fair to the model, none of this says Meituan does anything wrong with your data. It says Meituan has not published what it does. For a personal project that distinction does not matter. For a customer service automation pipeline carrying real tickets, it is the whole decision.
Who should use it
Good fit. Solo developers and small teams doing high-volume agentic coding inside a harness they control, where the tokens are the cost that matters and the data is not sensitive. Document conversion, scraping, codebase navigation, long refactors. The instruction-following and long-context coherence are real, and at $0.30 per million input the arithmetic is hard to argue with. If you are surveying open-source AI agents or picking a model for LLM optimization work, it belongs on the list.
Poor fit. Anyone who needs one-shot reasoning quality, a documented function-calling contract, a real 1M window, a card payment method, or a data-processing answer in writing. That covers most business buyers, and all of the best AI agent for customer service use cases I work on. Not because the model is bad, but because a raw model is the wrong unit of purchase for a support queue. The model is maybe 20% of the problem; retrieval, grounding, escalation rules, guardrails and testing are the other 80%, and none of those come in a checkpoint. That is the same conclusion I reach in build vs buy AI support, and the reason the cost per resolution number moves so much less with token price than people expect.
Try eesel
Reading a review like this is really an attempt to answer a different question: will this thing give my customers a wrong answer? Token price does not tell you. AI hallucination is not a line item on a pricing page.
That is the problem I work on. eesel simulates your AI agent against your own historical tickets before it ever replies to a live customer, so you see the answers it would have sent, on your real questions, with your actual knowledge base behind it. Every reply is logged, reviewable and reversible, and you set exactly which topics it is allowed to touch. It connects to your existing helpdesk AI stack in a few minutes, and model choice becomes eesel’s problem instead of yours. Free to try.

The verdict
LongCat 2.0 earns a recommendation for exactly one job: cheap, high-volume agentic coding in a harness you drive, on data you do not mind leaving your building. At the promo rate it is one of the best value-per-token options in its class, the MIT licence is the real thing, and the hands-on reports from people who ran billions of tokens through it are more positive than its benchmark row.
Everything blocking a wider recommendation is a documentation gap rather than a modelling one: a context number that does not match the config, a function-calling contract that was never published, a promo with no end date, and a data policy that does not exist. Meituan could close all four with a week of writing. Until it does, this is a great model to experiment with and a hard one to put in production.
If your actual goal is AI on a support queue rather than a coding agent, start from the best LLM for customer support instead, and treat the model as the last thing you pick, not the first.
Frequently Asked Questions
Is LongCat 2.0 any good?
What does this LongCat 2.0 review conclude about pricing?
Does LongCat 2.0 really have a 1M context window?
config.json on Hugging Face caps max_position_embeddings at 262,144, so the real serving context window is 256K. The 1M figure describes the training data. Output is separately capped at 131,072 tokens.How does LongCat 2.0 compare to Kimi K3 and Qwen3.8-Max?
Can I run LongCat 2.0 locally?
Is LongCat 2.0 safe for customer support use?
What is the best LongCat 2.0 alternative for support teams?
Does LongCat 2.0 support tool calling and MCP?
tools array in either endpoint schema, no MCP surface, and launch-day users reported a non-standard <longcat_tool_call> wrapper their harnesses could not parse. Tool use is possible through clients, but it is not a documented contract.
Article by
Alicia Kirana Utomo
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








