GLM-5.3 Flash: Specs, Price, Benchmarks and Open Weights
Updated 2026-10-11
GLM-5.3 Flash is an open-weight model from Z.ai (the company formerly called Zhipu AI), released on August 26, 2026. It is the first natively multimodal model in the GLM-5 series: a 320-billion-parameter mixture-of-experts model that activates about 18 billion parameters per token, takes text, images and video, and has a 1-million-token context window. Weights are MIT-licensed on Hugging Face and ModelScope, and the API costs $0.15 per million input tokens and $0.50 per million output tokens.
Below are the confirmed specs, how it compares with GLM-5.2 and GLM-5.3, and how to start using it. Figures come from Z.ai's launch as reported on August 26, 2026; check Z.ai's pricing page for current prices.
GLM-5.3 Flash specs
| GLM-5.3 Flash | |
|---|---|
| Released | August 26, 2026 |
| Developer | Z.ai (formerly Zhipu AI) |
| Architecture | Mixture-of-experts, hybrid sparse + linear attention |
| Parameters | 320B total, ~18B active per token |
| Context window | 1 million tokens |
| Inputs | Text, image, video (native multimodal) |
| License | MIT (open weights, FP8 and BF16) |
| API price | $0.15 / 1M input, $0.50 / 1M output, $0.03 / 1M cached input |
| DeepSWE v1.1 | 63.4 (GLM-5.2: 46.2) |
| Artificial Analysis Intelligence Index | 57 (at launch) |
Why people are searching for it
Three things made GLM-5.3 Flash stand out in late August:
- Price. At $0.15 / $0.50 per million tokens it is among the cheapest capable models; Z.ai describes it as roughly one-tenth of the cost of some competing options.
- Open weights with a permissive license. MIT lets companies use, modify and redistribute the model commercially.
- Chinese hardware. Z.ai deployed it entirely on domestically produced AI accelerators, with custom serving and quantization that it says tripled end-to-end performance versus a generic deployment – a notable point given US export limits on top NVIDIA chips.
The coding jump also drew attention: 63.4 on DeepSWE v1.1 versus 46.2 for GLM-5.2. Before launch it was tested anonymously; reporting at release identified the anonymous "Ox Alpha" model as GLM-5.3 Flash.
GLM-5.3 Flash vs GLM-5.3 vs GLM-5.2
- GLM-5.3 Flash – the fast, cheap, multimodal tier with open weights. Good default for agents, coding helpers and bulk work.
- GLM-5.3 – the larger sibling in the same generation, for harder reasoning where cost matters less. People searching "glm 5.3" often mean either one – check which model id your provider uses.
- GLM-5.2 – the previous generation; GLM-5.3 Flash beats it clearly on the coding benchmark above.
Z.ai has also announced a GLM-5.3-FlashX variant marketed for speed (around 200 tokens per second), aimed at latency-sensitive apps.
How to use GLM-5.3 Flash
- Z.ai API – sign up on Z.ai's developer platform and call the GLM-5.3 Flash model; the API follows the OpenAI chat format, so most SDKs work with a base-URL change.
- Gateways – several model routers list GLM-5.3 Flash; compare their price with Z.ai's own.
- Self-host – download the FP8 or BF16 weights from Hugging Face or ModelScope. With 320B total parameters you need multi-GPU server memory even though only ~18B are active per token; this is not a laptop model.
- In coding tools – many agent tools let you plug in any OpenAI-compatible endpoint. Point them at GLM-5.3 Flash for cheap first-pass coding work.
Clear instructions matter more with cheaper models. Our free ChatGPT prompt generator turns a rough task into a structured prompt that works with GLM as well.
FAQ
When was GLM-5.3 Flash released? August 26, 2026, by Z.ai.
Is GLM-5.3 Flash open source? The weights are released under the MIT license on Hugging Face and ModelScope, in FP8 and BF16.
How much does GLM-5.3 Flash cost? $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens on Z.ai's API at launch.
What is GLM-5.3 Flash's context window? 1 million tokens.
Can GLM-5.3 Flash read images and video? Yes – it is natively multimodal and accepts image and video input.
Can I run GLM-5.3 Flash locally? Only on serious hardware. The full model has 320B parameters, so it needs server-class GPU memory even though each token uses about 18B of them.