● GeneratorNest

GLM-5.3 Flash: Specs, Price, Benchmarks and Open Weights

Updated 2026-10-11

GLM-5.3 Flash is an open-weight model from Z.ai (the company formerly called Zhipu AI), released on August 26, 2026. It is the first natively multimodal model in the GLM-5 series: a 320-billion-parameter mixture-of-experts model that activates about 18 billion parameters per token, takes text, images and video, and has a 1-million-token context window. Weights are MIT-licensed on Hugging Face and ModelScope, and the API costs $0.15 per million input tokens and $0.50 per million output tokens.

Below are the confirmed specs, how it compares with GLM-5.2 and GLM-5.3, and how to start using it. Figures come from Z.ai's launch as reported on August 26, 2026; check Z.ai's pricing page for current prices.

GLM-5.3 Flash specs

GLM-5.3 Flash
Released August 26, 2026
Developer Z.ai (formerly Zhipu AI)
Architecture Mixture-of-experts, hybrid sparse + linear attention
Parameters 320B total, ~18B active per token
Context window 1 million tokens
Inputs Text, image, video (native multimodal)
License MIT (open weights, FP8 and BF16)
API price $0.15 / 1M input, $0.50 / 1M output, $0.03 / 1M cached input
DeepSWE v1.1 63.4 (GLM-5.2: 46.2)
Artificial Analysis Intelligence Index 57 (at launch)

Why people are searching for it

Three things made GLM-5.3 Flash stand out in late August:

  1. Price. At $0.15 / $0.50 per million tokens it is among the cheapest capable models; Z.ai describes it as roughly one-tenth of the cost of some competing options.
  2. Open weights with a permissive license. MIT lets companies use, modify and redistribute the model commercially.
  3. Chinese hardware. Z.ai deployed it entirely on domestically produced AI accelerators, with custom serving and quantization that it says tripled end-to-end performance versus a generic deployment – a notable point given US export limits on top NVIDIA chips.

The coding jump also drew attention: 63.4 on DeepSWE v1.1 versus 46.2 for GLM-5.2. Before launch it was tested anonymously; reporting at release identified the anonymous "Ox Alpha" model as GLM-5.3 Flash.

GLM-5.3 Flash vs GLM-5.3 vs GLM-5.2

  • GLM-5.3 Flash – the fast, cheap, multimodal tier with open weights. Good default for agents, coding helpers and bulk work.
  • GLM-5.3 – the larger sibling in the same generation, for harder reasoning where cost matters less. People searching "glm 5.3" often mean either one – check which model id your provider uses.
  • GLM-5.2 – the previous generation; GLM-5.3 Flash beats it clearly on the coding benchmark above.

Z.ai has also announced a GLM-5.3-FlashX variant marketed for speed (around 200 tokens per second), aimed at latency-sensitive apps.

How to use GLM-5.3 Flash

  1. Z.ai API – sign up on Z.ai's developer platform and call the GLM-5.3 Flash model; the API follows the OpenAI chat format, so most SDKs work with a base-URL change.
  2. Gateways – several model routers list GLM-5.3 Flash; compare their price with Z.ai's own.
  3. Self-host – download the FP8 or BF16 weights from Hugging Face or ModelScope. With 320B total parameters you need multi-GPU server memory even though only ~18B are active per token; this is not a laptop model.
  4. In coding tools – many agent tools let you plug in any OpenAI-compatible endpoint. Point them at GLM-5.3 Flash for cheap first-pass coding work.

Clear instructions matter more with cheaper models. Our free ChatGPT prompt generator turns a rough task into a structured prompt that works with GLM as well.

FAQ

When was GLM-5.3 Flash released? August 26, 2026, by Z.ai.

Is GLM-5.3 Flash open source? The weights are released under the MIT license on Hugging Face and ModelScope, in FP8 and BF16.

How much does GLM-5.3 Flash cost? $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens on Z.ai's API at launch.

What is GLM-5.3 Flash's context window? 1 million tokens.

Can GLM-5.3 Flash read images and video? Yes – it is natively multimodal and accepts image and video input.

Can I run GLM-5.3 Flash locally? Only on serious hardware. The full model has 320B parameters, so it needs server-class GPU memory even though each token uses about 18B of them.

Try it free

100% free – no sign-up needed.

Open the tool