GLM-5.3-Flash API Pricing
GLM-5.3-Flash is the first native multimodal model in the GLM-5 line, and it is priced like a small model without being one: 320B parameters with 18B active, at $0.15/M input and $0.50/M output — roughly a ninth of what GLM-5.3 costs. Z.ai ran a 50% opening discount on those rates until September 9, 2026; it ended on the date the vendor published, and the list figures on this page are now what you are charged.
Run the numbers.
Calculator pre-loaded with Z.ai's International rates. Until 24:00 on September 9, 2026 (UTC+8) an opening discount halved them; that window has closed, so these figures are both the list price and the bill. Thinking cannot be switched off on this model, so reasoning tokens land in the output count on every call.
Real-world presets.
Repo-wide feature build
Reading 150-page contracts
Support agent ticket triage
Research planning turn
Paste text. See tokens. See cost.
Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| GLM-5.3-Flash Current | $0.15 cache $0.03 | $0.50 | $0.0875 list price | 1M | Multimodal frontier work at a light-tier rate |
| GLM-4.7-Flash | Free cache Free | Free | Free free tier | 200K | Free, and a generation behind |
| GLM-4.7-FlashX | $0.07 cache $0.01 | $0.40 | $0.0511 cheaper | 200K | Cheapest paid GLM, no multimodal |
| GLM-4.7 | $0.60 cache $0.11 | $2.20 | $0.358 pricier | 200K | Previous mid tier, fifth the context |
| GLM-5.3 | $1.40 cache $0.26 | $4.40 | $0.78 pricier | 1M | The full-size flagship, text only |
| Qwen3.8-Flash | $0.15 | $0.47 | $0.176 pricier | 1M | Same list input, no published cache rate |
| Gemini 3.7 Flash | $0.75 cache $0.075 | $3.75 | $0.481 pricier | 1M | The Western flash tier, five times the rate |
Frequently asked.
GLM-5.3-Flash pricing questions, including what happened to the opening discount.
Q · 01 How much does GLM-5.3-Flash cost? +
$0.15/M input, $0.03/M on a cache hit and $0.50/M output. Those were always the list rates; a 50% opening discount charged half of each — $0.075, $0.015 and $0.25 — until 24:00 on September 9, 2026 (UTC+8), and Z.ai let it expire on the published date rather than extending it. Because this catalogue stored the list throughout, no figure on this page had to change when the discount lapsed.Q · 02 Why did this page quote the list price while the discount was running? +
Q · 03 How does it compare with GLM-5.3? +
$1.4/M input and $4.4/M output — 9.3x and 8.8x this model's list rate. On the agentic blend the gap is 8.9x. The Flash is also the more capable model in one respect the flagship cannot match: it is natively multimodal, where GLM-5.3 is text in and text out.Q · 04 Can thinking be turned off? +
thinking.type only supports enabled on this model, so reasoning tokens are billed as output on every call. Budget the output side more generously than a non-reasoning model of the same price would need — the vendor recommends reasoning_effort: max and setting thinking.clear_thinking to false.Q · 05 What does it accept as input? +
Q · 06 What about the free cached-input storage? +
$0.03/M cache-hit rate above is unaffected by it.