● GeneratorNest

Qwen 3.8-Max: Specs, Rankings and How to Use It

Updated 2026-10-11

Qwen 3.8-Max is Alibaba's flagship model in the Qwen 3.8 generation, announced on August 3, 2026. It is a sparse mixture-of-experts model with 2.4 trillion total parameters, of which about 95 billion are active per token, and it supports a context window of up to 1 million tokens. At launch Alibaba reported it ranked 5th on LMArena's Text Arena, 2nd on Vision Arena and 4th on Frontend Code Arena. Developers use it through Alibaba Cloud's Model Studio API; consumers can try it in Alibaba's "Qwen Work" office-agent platform.

Here is what Alibaba has confirmed about Qwen 3.8-Max, how it fits with the other Qwen 3.8 models people search for, and how to start using it. Arena rankings move every week, so treat the launch positions as a snapshot.

Qwen 3.8-Max specs

Qwen 3.8-Max
Announced August 3, 2026
Developer Alibaba (Qwen team)
Architecture Sparse mixture-of-experts + hybrid attention, built on Qwen 3.5
Total parameters ~2.4 trillion
Active parameters ~95 billion per token
Context window Up to 1 million tokens
Inputs Text, images, video (natively multimodal)
Launch rankings Text Arena #5, Vision Arena #2, Frontend Code Arena #4
Access Alibaba Cloud Model Studio API; Qwen Work

The point of the mixture-of-experts design is cost: only a small slice of the 2.4T parameters runs for each token, so inference is much cheaper and faster than a dense model of the same size.

What Alibaba says it can do

  • Long autonomous coding. In an internal test, Alibaba says Qwen 3.8-Max ran a real software project for 16 days on its own. Asked to build a self-improving agent framework from scratch, it produced "oh-my-cli", which Alibaba open-sourced on GitHub.
  • Office and professional work. Alibaba lists app design, legal document review, sports data analysis, financial research and 3D architectural modelling as target workloads.
  • Vision and video. It can process documents of 100+ pages, a whole TV series or up to 100 hours of livestream and turn them into searchable knowledge. Examples include editing raw clips into a vlog, rebuilding a front-end project from one UI screenshot and turning a 2D floor plan into a 3D interior render.
  • RecreationBench. Alibaba introduced this benchmark for rebuilding an app from scratch only by using it – no internet, no source code. These are Alibaba's own results; independent replications were limited at the time of writing.

Qwen 3.8, Qwen 3.8-Max and the other 3.8 models

"Qwen 3.8" searches mix several models:

  • Qwen 3.8-Max – the flagship described above (API first; Alibaba said at launch that weights would follow).
  • Qwen3.8-27B – a much smaller model released alongside Max, the one most people can run locally.
  • Qwen3.8-Flash / Flash-Next – the fast, cheaper tier. Flash-Next is the 125B mixture-of-experts model behind the "run it on a 12 GB GPU" story – see our guide to Strata and Qwen 3.8.

If you want the strongest results and don't mind an API, use Max. If you want to run something on your own hardware, look at 27B or Flash-Next.

How to use Qwen 3.8-Max

  1. API: create an Alibaba Cloud account, open Model Studio (international) and select Qwen 3.8-Max. The API is OpenAI-compatible, so most SDKs and agent tools work with a base-URL change.
  2. Via gateways: model routers such as OpenRouter list Qwen models; check that the exact "3.8-Max" variant is offered and compare the price with Alibaba's own.
  3. In a chat app: Qwen Work and Qwen Chat let you try it without code.
  4. Prompt for long tasks: give it a clear goal, constraints and a definition of "done" – that is where agentic models gain the most. Our free ChatGPT prompt generator works for Qwen prompts too.

FAQ

When was Qwen 3.8-Max released? Alibaba announced it on August 3, 2026.

How many parameters does Qwen 3.8-Max have? About 2.4 trillion in total, with roughly 95 billion active per token.

Is Qwen 3.8-Max open source? At launch it was available through the API, and Alibaba said the model weights would be published shortly after. Check the Qwen pages on Hugging Face or ModelScope for the current status and license.

What is Qwen 3.8-Max's context length? Up to 1 million tokens.

Is Qwen 3.8-Max better than GPT or Claude? At launch it ranked 5th on Text Arena, so a few frontier models were ahead on general chat, while it ranked 2nd for vision. Test it on your own tasks – rankings shift weekly.

Can I run Qwen 3.8-Max locally? Not realistically on consumer hardware – even with only ~95B active parameters, the full 2.4T weights need datacenter-scale memory. Use Qwen3.8-27B or Flash-Next for local use.

Try it free

100% free – no sign-up needed.

Open the tool