← Back to Blog
AI News·8 min read

Qwen3.8-Max Review: I Tested Alibaba's 2.4T Model

Listen to this article

Alibaba dropped Qwen3.8-Max-Preview on 19 July 2026 at WAIC in Shanghai, a few days after Moonshot launched Kimi K3, and the announcement came with a bold claim: "second only to Fable 5". No benchmark table. No model card. No technical report. Just a tweet, a 2.4 trillion parameter figure, and a preview endpoint. So this Qwen3.8-Max review is me doing what Alibaba didn't: actually testing the thing.

Quick naming note before we start. The official name is Qwen3.8-Max-Preview (product ID qwen3.8-max-preview). Everyone's calling it "Qwen 3.8 Max", including me in casual conversation, but keep in mind this is a paid preview available through Alibaba's Token Plan, not a finished release. Open weights are promised "soon", with no date and no licence. We'll see.

This month has been absolutely insane for AI releases. I still owe you reviews of ChatGPT 5.6 and Kimi K3 (I've covered Grok 4.5 and Claude Fable 5 already), but I'm reviewing Qwen first for a simple reason: Qwen has quietly become the most important model family for local AI. If you run agents locally, chances are you're running Qwen. So when they ship a new flagship, I want to know what's coming down the pipeline to the smaller models I'll eventually run on my own hardware.

What We Actually Know (and Don't)

Here's the honest state of play as of today.

Confirmed: 2.4T total parameters (vendor figure), sparse MoE, multimodal (text, images, video, documents). It's the Qwen team's first multimodal model above 1T parameters. The endpoints speak both OpenAI and Anthropic protocols, so it drops straight into Claude Code, Cursor, OpenCode and friends without changing your harness.

Not confirmed: basically everything else. Context window (the 1M figure floating around is community lore, not an official spec), active parameter count, knowledge cutoff, benchmarks. Every benchmark number you'll see circulating "for" Qwen3.8 is actually a Qwen3.7-Max number. The only independent evaluation I've found is Trilogy AI's StackPerf run, a single blind head-to-head where Qwen3.8-Max-Preview scored 80 against Kimi K3's 83. One run, one task, so take it as a data point, not a verdict.

Qwen3.8-Max-Preview at a glance:

Released 19 July 2026 (preview)
Parameters 2.4T total (vendor-reported), MoE, active count undisclosed
Access Token Plan, Qoder, QoderWork (credits only, no per-token API price)
Preview discount 1/10th standard rate, 1/50th overnight
Open weights Promised "soon", no date, no licence
Benchmarks None published by Alibaba

My Test Suite: Same 4 Prompts, Every Model

If you've read my other reviews, you know the drill. I run the same four tests with the same prompts on every model so I can compare directly: a Sydney coffee roaster landing page, a pop culture themed clothing store, a Texas Hold'em poker simulation written in Go (six AI players with distinct personalities, realistic betting, 1000 hands, statistics at the end), and a full audit of this website. Same methodology as my Grok 4.5 and Gemini 3.5 Flash reviews.

The Results

Coffee Roaster Website

This took a looong time. The model sat there thinking for several minutes before doing anything at all, and it stayed slow throughout. Over 30 minutes total, the longest this test has ever taken.

But here's the thing: the result might be the best I've ever gotten from this prompt. It went into obsessive detail after the first HTML hit the disk, testing repeatedly with Playwright and adding little touches each pass. I got things I never asked for: a shopping cart, a wholesale section with registration, an Instagram link, coffee sorting and search by roast level. In a different context that scope creep could be a problem. Here, it made the site feel real. Visually it was tasteful too. Nice fonts, nice colours, good animations, solid layout.

Pop Culture E-commerce Store

Even slower than the coffee site. At this point I started wondering whether the model is just fundamentally slow, or whether Alibaba simply didn't have enough servers ready for launch day (I was testing hours after release, to be fair).

The obsessiveness continued. It checked buttons, menus, animations and aspect ratios with Playwright like it was getting paid per test run. The result was great: polished, detailed, good typography. I still slightly prefer the style Gemini 3.5 Flash gave me on this prompt, but that's taste. If the speed improves, this is a serious frontend model.

Poker Simulation (Go)

One hour and twenty minutes. I made coffee. I answered emails. I briefly considered learning to whittle.

But it one-shotted it. Only two other models have managed that so far, Claude Fable 5 and Grok 4.5, which makes Qwen3.8-Max the third member of a very small club. What's interesting is that the core simulation seemed done after about 20 minutes. The remaining hour went into fine-tuning the six player archetypes (the tight rock, the crazy maniac, and so on) until the betting behaviour looked genuinely realistic. And it did. The final simulation output showed clearly distinct, believable playing styles across 1000 hands. The statistics were good, though not as detailed as what Grok 4.5 produced on the same test.

Website Audit

35 minutes, the slowest audit in this test series. It didn't find any new errors on thomas-wiegold.com, but it flagged potential pitfalls I'd genuinely never thought about. And honestly, an audit is the one test where slow and thorough is exactly what you want. My tip if you care about this stuff: run the same audit through Fable, ChatGPT 5.6, Kimi K3 and Qwen3.8-Max, then consolidate the results into one file to work through. The overlap tells you what's real, the differences tell you what's interesting.

The Speed Problem

Let's not dance around it. In every single test, this was the slowest model I've used. The open question is why. It could be the model itself being extremely verbose in its thinking. It could be day-one server strain, since I tested hours after launch. Alibaba calls the preview "continuously evolving", so this may well improve.

The trade-off is real though. For rapid iteration, the speed is genuinely annoying. For important one-shot builds where you want maximum thoroughness, watching it Playwright-test every button for an hour starts to feel less like slowness and more like diligence. I'll update this review if the speed changes.

Pricing: The Token Plan Deal

I subscribed to the Token Plan Standard tier at USD $18/month for this review. All four tests combined used 6% of my weekly credit quota. Six percent. For a model this capable, that's a genuinely good deal, and the plan also gives you access to other models plus image generation.

The catch is the preview discount doing heavy lifting: Qwen3.8-Max-Preview currently runs at 1/10th of standard credit rates, dropping to 1/50th overnight (22:00 to 08:00 UTC+8, so basically midnight to 10am here in Sydney). When the preview pricing ends, redo your maths. And remember Alibaba's own docs say the preview endpoint may change or be replaced. Test on it, don't build production on it.

Verdict: Who Should Use Qwen3.8-Max?

Qwen3.8-Max is very good and very slow. If speed doesn't matter for your task, it belongs in the top tier right now, alongside Fable 5, ChatGPT 5.6 Sol, Grok 4.5 and Kimi K3. Honestly, they're all great, and they're all close. The frontier has become crowded, and the real differentiators now are price, speed, availability and vibe.

For me, today, Grok 4.5 stays my daily driver purely because of speed and efficiency. That could change next month. It changes depending on the task, too. And if you're doing something genuinely important, run it through two or three of these models and compare. That's a luxury we simply didn't have a year ago.

Which brings me to Anthropic. In my opinion they're the big loser of this month. After telling everyone Fable is too good and too expensive to include permanently in normal subscriptions, they've watched competitors ship models that are (probably) roughly as good, much cheaper, more available, and less restricted. The competition didn't out-build them, it out-shipped them. No wonder people are speculating Opus 5 is coming soon. For a while I was genuinely worried the top models would end up locked behind enterprise contracts. Not anymore, and I'm very happy about that.

One caveat worth knowing before you evaluate Qwen for business use: Simon Willison and others have documented that Qwen models carry baked-in CCP-aligned guardrails on politically sensitive topics. For coding and agent work it's largely irrelevant, but if you're a regulated business doing anything content-adjacent, know what you're deploying. I wrote more about choosing AI tools sensibly in my piece on AI agents for small business.

And the thing I'm actually most excited about? The small models. Qwen is the open-weight family powering half the local agent setups out there, including my own experiments with Hermes agent. If the promised open weights land, and if the 3.8 improvements trickle down to sizes I can run on my Mac, that's the real story. I think local AI is going to be very, very good within a few years, especially for businesses like medical practices and law firms where privacy and legal constraints make cloud AI complicated. When the small Qwen3.8 models drop, I'll be testing them the same week. Watch this space.

Thomas Wiegold

AI Solutions Developer & Full-Stack Engineer with 15+ years of experience building custom AI systems, chatbots, and modern web applications. Based in Sydney, Australia.

Ready to Transform Your Business?

Let's discuss how AI solutions and modern web development can help your business grow.

Get in Touch