
The 60-second table
Everything below is from each vendor's own published pages as of August 3, 2026. Where the two vendors measure the same thing differently, I have said so rather than smoothing it over.
| Qwen 3.8 Max | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | |
|---|---|---|---|---|
| Model ID | qwen3.8-max | gpt-5.6-sol | gpt-5.6-terra | gpt-5.6-luna |
| Released | Aug 2, 2026 | Jul 9, 2026 | Jul 9, 2026 | Jul 9, 2026 |
| Input / 1M | $2.00 | $5.00 | $2.00 | $0.20 |
| Output / 1M | $6.00 | $30.00 | $12.00 | $1.20 |
| Cached input / 1M | $0.25 | $0.50 | $0.20 | $0.02 |
| Long-context surcharge | None published | 2x in, 1.5x out above 272K | Same | Same |
| Context window | 1,000,000 | 1,050,000 | 1,050,000 | 1,050,000 |
| Max output | 131,072 | 128,000 | 128,000 | 128,000 |
| Max reasoning tokens | 262,144 | Not published | Not published | Not published |
| Input modalities | Text, image, video | Text, image | Text, image | Text, image |
| Parameters | 2.4T total, 95B active | Not published | Not published | Not published |
| Knowledge cutoff | Not published | Feb 16, 2026 | Feb 16, 2026 | Feb 16, 2026 |
| Model card | None | Yes | Yes | Yes |
| Open weights | Promised, not shipped | No | No | No |
| Rate limit (top tier) | 15K RPM / 2M TPM | 15K RPM / 40M TPM | Same as Sol | 30K RPM / 180M TPM |
| Available in consumer chat | Qwen Studio | ChatGPT | Codex and Work only | Codex and Work only |
| Reasoning control | reasoning_effort: low / medium / xhigh | Reasoning effort incl. max | Same | Same |
Two rows deserve a second look. Qwen is the only one of the four that takes video as input, which is a real capability gap and not a spec-sheet flourish. And Qwen is the only one with no published knowledge cutoff, no model card and no system card, which is a real evidence gap.
What Qwen 3.8 Max actually is
Qwen is Alibaba Cloud's model family, and Max is the proprietary flagship line that sits above the cheaper Plus and Turbo tiers. Qwen 3.8 Max is the newest one, and it shipped in two acts.
If you want the wider family first, my Qwen review covers the older tiers, and the standalone Qwen 3.8 Max review goes deeper on this model alone.
Act one was July 19 at the World AI Conference in Shanghai: a preview endpoint, two posts on X, and a spoken frontier-ranking claim with no chart. Act two was August 2, when Alibaba published a 5,000-word launch post with everything the preview was missing.
The architecture, as Alibaba states it: 2.4 trillion total parameters with 95 billion active, a sparse Mixture-of-Experts model built on the Qwen 3.5 foundation. That 95B active figure is the one that matters and the one the preview withheld. Total parameters set the headline; active parameters set your latency and your bill.
The multimodal claim is the genuinely new thing for the Max line. Alibaba's model page lists image, text and video in, and the launch post puts numbers on it: financial reports and PDFs past 200 pages, and video "longer than 100 hours" organised into what Alibaba calls a video memory graph. Worth flagging honestly, though: both harness configs Alibaba publishes in that same post declare text and image only. Video looks like a native DashScope path, not something you get through the OpenAI-compatible endpoint.

Reasoning is controlled with reasoning_effort, which takes low, medium and xhigh, and xhigh is the default. Alibaba also enables preserve_thinking by default "for best out-of-the-box experience," which is a polite way of saying the model is tuned to think hard before it answers. Hold that thought, because it shows up later as a cost problem.
What GPT-5.6 actually is
GPT-5.6 is not one model. It is three capability tiers that went generally available on July 9: Sol as the flagship, Terra as the balanced middle, and Luna as the fast, cheap one. I broke each of them down in a full GPT-5.6 review. All three share the same 1,050,000-token context window, the same 128,000-token output ceiling, and the same February 16, 2026 knowledge cutoff.
Then on July 30, OpenAI did the thing that quietly reset this whole comparison. It cut prices on two of the three tiers:
"Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna. Sol pricing remains unchanged."
That is 20% off Terra and 80% off Luna. Sol held at $5 and $30. If you read a GPT-5.6 pricing comparison written in mid-July, it is wrong by up to 5x on Luna.
The same update renamed Priority Processing to Fast mode, which doubles the price for up to 2.5x the speed on Sol. And OpenAI added something Qwen has not: a long-context surcharge. Prompts over 272K input tokens are priced at 2x input and 1.5x output for the whole request, not just the overflow. Crossing that line turns a $5/$30 Sol call into a $10/$45 one. So Sol's real million-token window is a million tokens you pay double for.
Benchmarks: two tables, no referee
Here is where most head-to-heads go wrong. Both labs now publish a chart. Neither chart was run by anyone independent, and the two cannot be stacked on top of each other.

Start with what Alibaba published, because it is unusually direct: its chart carries GPT-5.6 Sol (max) as a comparison column. Straight from that chart:
| Benchmark | Qwen 3.8 Max | GPT-5.6 Sol (max) | Winner |
|---|---|---|---|
| SWE-Pro | 67.7 | 64.6 | Qwen |
| TerminalBench-2.1 | 86.6 | 88.8 | Sol |
| PaperBench | 93.0 | 90.5 | Qwen |
| FrontierSWE | 73.5 | 88.8 | Sol, by 15.3 |
| CoWorkBench | 74.8 | 71.5 | Qwen |
| JobBench | 53.4 | 45.4 | Qwen |
| Agents' Last Exam | 52.4 | 53.6 | Sol |
| CharXiv (RQ) | 93.5 | 89.1 | Qwen |
| ERQA | 77.8 | 70.0 | Qwen |
| PerceptionBench | 63.5 | 59.7 | Qwen |
| LVBench | 81.8 | 78.8 | Qwen |
| Vision2Web | 69.0 | 62.1 | Qwen |
| OSWorld-Verified | 86.1 | 83.2 | Qwen |
Thirteen of sixteen to Qwen, and every vision row is a clean Qwen win. That is a real result and I am not going to talk it down.
But read the footnotes, which almost nobody does. Alibaba discloses that six of the coding benchmarks are its own in-house evals (QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, CoWorkBench, RecreationBench), that competitor scores were pulled from mixed harnesses rather than run under identical conditions, and that several rows are graded by rival models: gemini-3.1-pro-preview judges PLawBench, Claude Opus 4.6 judges PaperBench. There is also a line admitting Fable 5's results "may involve fallbacks."
Now the third parties, which split the verdict cleanly down the middle. On Artificial Analysis, whose Intelligence Index v4.1 is an automated composite of nine evals, GPT-5.6 Sol scores 59 and ranks third overall behind Claude Opus 5 and Claude Fable 5, with Terra at 55 and Luna at 51.
On LMArena, which is human preference voting, qwen3.8-max sits at #5 on Text with 1496 and #4 on WebDev with 1668, both ahead of gpt-5.6-sol-xhigh at #15 and #6. The same automated-versus-human split shows up in my GPT-5.6 vs Claude comparison.
So: automated composite says Sol, human preference says Qwen. Anyone citing only one of the two is selling you something.
The community read is blunter. From the Hacker News thread on the preview:
"The few tests I ran were by no means comprehensive, but while kimi felt like the real deal qwen seems a bit of a benchmark princess."
The price comparison everyone gets backwards
The story that wrote itself in July was that a Chinese lab had undercut OpenAI. At $2 and $6 against Sol's $5 and $30, that is true. Against the rest of the family, it is not.

Qwen 3.8 Max matches Terra's input price to the cent and halves its output price. That is a genuine win. But Luna undercuts Qwen by 10x on input and 5x on output, and it did not exist at that price when the "Chinese models are cheaper" framing set.
There is a second-order problem too, and it is the one that decides real bills. Qwen ships with xhigh reasoning on by default and preserve_thinking enabled. Verbose models cost more than their sticker rate suggests, because you pay for every reasoning token. Artificial Analysis measured Luna burning 130M output tokens to run its index against a 62M median, so this cuts both ways. The point is that rate per token and cost per task are different numbers, and only one of them is on the pricing page.
Plug your own volumes in:
One more wrinkle if you are looking at the cheap monthly subscription instead of the API. The qwen3.8-max-preview endpoint was never sold per token at all: it was only available through Alibaba's Token Plan, at $6, $18 or $68 a month on the Personal tiers. Those plans run on credits with fixed 5-hour and 7-day windows, unused credits do not roll over, and hitting either cap pauses you until the window resets. Alibaba also declines to publish the credits-per-token coefficient, so you cannot compute your cost in advance. The headline preview deal, 10% of normal rates stacking with an extra 80% off between 22:00 and 08:00, works out to 2% of standard burn, applies to Personal plans only, and dies with the preview. The GA model gets a flat 50% night discount instead.
What actually decides it in practice
Benchmarks measure capability. Production measures capability divided by patience. Every hands-on report I found lands on the same complaint, and it is not about intelligence.
A developer who ran the same rich design conversion through both models on release day:
"Same prompt for both for the conversion. I used OpenCode for the qwen version, but I encountered a significant amount of errors / timeouts while it was running. Claude finished in around 16 min, but I spent close to 2 hours shepherding the Qwen build."
Another, after testing four frontier models side by side over several days:
"Don't trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA."
That same developer posted the token accounting for one web-app task, which is the most useful cost data anyone has published: Kimi K3 at $5.50, Qwen 3.8 Max at $6.30, Fable 5 at $30. Qwen's cheap rate is real, and it still needed 18 million input tokens to Kimi's 9.5 million for the same job.
Verdicts on the model itself are split, which is worth saying plainly rather than averaging away. One user cancelled an Anthropic subscription over it. Another called it "disastrously bad," describing a model that "gets stuck into long second-guessing loops with no progress." Both were testing the same preview build in the same fortnight. If that spread bothers you, the safer shortlists are my Qwen alternatives and GPT-5.6 alternatives roundups, plus the Kimi K3 review for the third contender in this bracket.
And a few things simply are not published. There is no knowledge cutoff for Qwen 3.8 Max anywhere, no model card, no system card, no safety evaluation. The open weights promised for "next week" from August 2 have no license, no date and no repo. If your procurement process asks for any of that, the answer today is that it does not exist.
Where a frontier model stops and support work begins
This is the part I care about most, because it is where I watch teams lose a quarter.

Neither of these models knows your refund policy. Neither has read your last 7,000 tickets. Neither has any concept that the customer emailing you is on an enterprise plan and has already been escalated twice. A benchmark score is not a support agent, and the gap between the two is where most AI support automation projects quietly die.
The same gap is why picking an AI helpdesk is a different exercise from picking a model, and why ticket automation lives or dies on retrieval rather than raw reasoning.
We have spent years running AI on live support queues, and the specific failure mode is worse than "the AI does not know." It is confident invention. We have had paying customers whose bots fabricated answers to real customers when knowledge-base retrieval came back empty: one solar company's bot invented subscription terms that did not exist, and another sent a customer "Oxygen," from the periodic table, as an answer. Nothing in a 2.4-trillion-parameter model prevents that. A bigger model just says the wrong thing more fluently.
The control that actually fixes it came up in a call with a support lead running 7,000 tickets a month, and it has nothing to do with model choice:
"The AI will never be able to answer 100% of the questions, but if it tries and just answers 'sorry I don't know this,' I cannot go and check all my 7,000 tickets to see if the AI actually made a good answer. I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."
a CX lead at a 7,000-ticket-a-month DTC brand
That is a product decision, not a model decision. It is why eesel AI runs a simulation on historical tickets before anything goes live, so you see the resolution rate and the actual replies on real past conversations first.
It is also why answers route by confidence, drafting for a human or escalating rather than guessing. Running it that way, Gridwise hit 73% tier-1 resolution in its first month.
The strategic read on Qwen 3.8 Max versus GPT-5.6 is therefore a little anticlimactic: build so you can swap the engine. The model layer gets better and cheaper every few weeks. Kimi K3 landed days before Qwen, Luna's price fell 80% in a single afternoon, and whatever tops the chart in September is not on it today. A team that hard-wired itself to one API in July is already re-plumbing.
That is also the honest answer to "should I run Gemini, Grok or Mistral instead." Probably, eventually, for some task. Which is the argument for not caring very much about this month's winner.
So which one should you pick?
If you are shipping product and want one answer:
- Vision, video or document-heavy work: Qwen 3.8 Max, comfortably. Every vision row on Alibaba's chart is a Qwen win, and it is the only one of the four that takes video at all.
- Hard software engineering: GPT-5.6 Sol. FrontierSWE at 88.8 against 73.5 is not a rounding error, and it is Alibaba's own number.
- High-volume, latency-sensitive, cost-sensitive: GPT-5.6 Luna, and it is not close on price. Watch the verbosity.
- A general workhorse where you want the best rate-per-capability: Terra and Qwen 3.8 Max are the real fight. Qwen is half the output price; Terra is faster and comes with a model card, a knowledge cutoff and a stable endpoint.
- Anything customer-facing: neither, on its own. Pick the layer first and let it pick the model. The maths is in my agent cost breakdown, and the shortlist is in best AI agents.
The thing I would not do is treat the August 2 chart as settled. Six of those coding benchmarks are Alibaba's own, no third party has rerun any of it, and the community verdict on the same build ranges from "cancelled my Anthropic subscription" to "disastrously bad." Give it a month and independent numbers.
Try eesel
If you landed here because you want a model to resolve tickets rather than win charts, that is the whole point of eesel AI. It plugs into Zendesk, Freshdesk and 100+ other tools, learns from your help center and past tickets, and starts drafting in minutes.

The differentiator is the trust ramp, not the engine. Simulate on your real ticket history, read the numbers, start in draft mode, and go autonomous only when you are happy. You inherit every frontier gain, whether that turns out to be Qwen, GPT-5.6 or whatever ships next month, without re-plumbing anything. Try eesel free, no credit card needed.
Frequently Asked Questions
Is Qwen 3.8 Max cheaper than GPT-5.6?
Which is better for coding, Qwen 3.8 Max or GPT-5.6 Sol?
What is the context window on Qwen 3.8 Max vs GPT-5.6?
Is Qwen 3.8 Max open source?
Can I use Qwen 3.8 Max or GPT-5.6 for customer support?

Article by
Rama Adi Nugraha
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








