Last verified
FLASH-TIER GLM-5NATIVE MULTIMODAL320B / 18B ACTIVE1M CONTEXTTHINKING ALWAYS ON

GLM-5.3-Flash API Pricing

GLM-5.3-Flash is the first native multimodal model in the GLM-5 line, and it is priced like a small model without being one: 320B parameters with 18B active, at $0.15/M input and $0.50/M output — roughly a ninth of what GLM-5.3 costs. Z.ai ran a 50% opening discount on those rates until September 9, 2026; it ended on the date the vendor published, and the list figures on this page are now what you are charged.

Input - per 1M tokens
$0.15/M
Base token price - a ninth of GLM-5.3 vs GLM-5.3
Output - per 1M tokens
$0.50/M
Output tokens - reasoning cannot be switched off thinking billed here
Cached input - per 1M tokens
$0.03/M
Cache hit, 5x cheaper than base -80%
Effective - agentic blend
$0.0875/M
92/8 split - 82% cache hit rate
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with Z.ai's International rates. Until 24:00 on September 9, 2026 (UTC+8) an opening discount halved them; that window has closed, so these figures are both the list price and the bill. Thinking cannot be switched off on this model, so reasoning tokens land in the output count on every call.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 450
Words 76
Tokens (estimated) 86 tokens
Cost as input · uncached $0.00001 USD
Cost as output · uncached $0.00004 USD
Cost as cached input $0.000003 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-5.3-Flash Current $0.15 cache $0.03 $0.50 $0.0875 list price 1M Multimodal frontier work at a light-tier rate
GLM-4.7-Flash Free cache Free Free Free free tier 200K Free, and a generation behind
GLM-4.7-FlashX $0.07 cache $0.01 $0.40 $0.0511 cheaper 200K Cheapest paid GLM, no multimodal
GLM-4.7 $0.60 cache $0.11 $2.20 $0.358 pricier 200K Previous mid tier, fifth the context
GLM-5.3 $1.40 cache $0.26 $4.40 $0.78 pricier 1M The full-size flagship, text only
Qwen3.8-Flash $0.15 $0.47 $0.176 pricier 1M Same list input, no published cache rate
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 pricier 1M The Western flash tier, five times the rate
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

GLM-5.3-Flash pricing questions, including what happened to the opening discount.

Q · 01 How much does GLM-5.3-Flash cost? +
$0.15/M input, $0.03/M on a cache hit and $0.50/M output. Those were always the list rates; a 50% opening discount charged half of each — $0.075, $0.015 and $0.25 — until 24:00 on September 9, 2026 (UTC+8), and Z.ai let it expire on the published date rather than extending it. Because this catalogue stored the list throughout, no figure on this page had to change when the discount lapsed.
Q · 02 Why did this page quote the list price while the discount was running? +
Because a promotion that ends quietly would otherwise turn every stored figure false overnight, and because comparisons across 300 rows only hold if they are all on the same basis. Chinese vendors run these discounts constantly and often without an end date; recording the list keeps the catalogue consistent and errs toward over-estimating your bill rather than under-estimating it. September 9 is what that policy is for: the discount ended and not one number on this page had to move.
Q · 03 How does it compare with GLM-5.3? +
GLM-5.3 lists at $1.4/M input and $4.4/M output — 9.3x and 8.8x this model's list rate. On the agentic blend the gap is 8.9x. The Flash is also the more capable model in one respect the flagship cannot match: it is natively multimodal, where GLM-5.3 is text in and text out.
Q · 04 Can thinking be turned off? +
No. Z.ai documents that thinking.type only supports enabled on this model, so reasoning tokens are billed as output on every call. Budget the output side more generously than a non-reasoning model of the same price would need — the vendor recommends reasoning_effort: max and setting thinking.clear_thinking to false.
Q · 05 What does it accept as input? +
Images, video and files, natively — Z.ai calls it "the first native multimodal model in the GLM-5 series" and the visual capability is built into the coding loop rather than bolted on, so the model can look at rendered output and iterate. Text parameters match GLM-5.3, including the 1M-token context window.
Q · 06 What about the free cached-input storage? +
Z.ai marks cached input storage as limited-time free across the whole GLM line. That is a separate charge from the token rates — a fee for holding the cache, not for reading it — and this catalogue does not model it. The $0.03/M cache-hit rate above is unaffected by it.