Last verified
DEPRECATEDLAST LISTED AUG 24200K CONTEXT128K OUTPUTTEXT + CODE

GLM-5 Turbo API Pricing

Off the rate card: GLM-5 Turbo was Zhipu's latency-optimized GLM-5 variant, and Z.ai no longer lists or prices it. Its row was on the International card on August 24, 2026 and gone by August 29 — the same week GLM-5.3-Flash was added — with no retirement notice published and no entry left in the Z.ai docs index. The figures below are frozen at the last published values: $1.20/M input, $4/M output, $0.24/M on a cache hit. We hold no Z.ai key, so whether the endpoint still answers is not something this page can claim; what it can show is that the vendor stopped selling it. For a current model at this tier see GLM-5.2, or GLM-5.3-Flash at a ninth of the price.

Input - per 1M tokens
$1.20/M
Last listed 2026-08-24 frozen
Output - per 1M tokens
$4.00/M
Context 200K, as published frozen
Cached input - per 1M tokens
$0.24/M
Cache hit, 5x cheaper than base -80%
Effective - agentic blend
$0.70/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the last rates Z.ai published for GLM-5 Turbo. They are a historical record, not a quote: the model is off the card, so use this to read old invoices or to size the gap against a model you can still buy.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 511
Words 87
Tokens (estimated) 97 tokens
Cost as input · uncached $0.00012 USD
Cost as output · uncached $0.00039 USD
Cost as cached input $0.00002 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-5 Turbo Current $1.20 cache $0.24 $4.00 $0.70 agentic 92/8 200K Latency-optimized GLM-5
GLM-5.1 $1.40 cache $0.26 $4.40 $0.78 pricier 200K Flagship agentic coding
GLM-5 $1.00 cache $0.20 $3.20 $0.572 cheaper 200K Cheaper GLM-5 family base
GLM-4.7 $0.60 cache $0.11 $2.20 $0.358 cheaper 200K Mid-tier GLM agents
GLM-4.7 FlashX $0.07 cache $0.01 $0.40 $0.0511 cheaper 200K Ultra-cheap GLM traffic
GLM-4.7 Flash Free cache Free Free Free cheaper 200K Free registered-user tier
DeepSeek V4 Pro $1.32 cache $0.044 $3.96 $0.569 cheaper 1M DeepSeek frontier discount tier
Qwen3 Coder Plus $1.00 $5.00 $1.32 pricier 1M Qwen code flagship

Frequently asked.

Practical pricing questions, separated from calculator assumptions and regional tiers.

Q · 01 What is GLM-5 Turbo priced at? +
Nothing, any more — Z.ai has taken it off the price card. The last published rates were $1.20/M input, $4/M output and $0.24/M on a cache hit, and those are the figures frozen on this page. The row was still on the card on August 24, 2026 and gone by August 29.
Q · 02 How is the effective price calculated? +
AI//COST uses the same 92/8 agentic blend everywhere. With an 82% cache hit rate, GLM-5 Turbo's effective blended cost is $0.7/M.
Q · 03 Is prompt caching priced separately? +
It was: the vendor table put cached input at $0.24/M against $1.2/M fresh, a fifth of the rate. Z.ai still applies that structure to the models it does sell — see the GLM lineup — and still shows free cached input storage across the range.
Q · 04 Are regional prices different? +
Z.AI publishes the official developer pricing page in USD. Chinese BigModel pages may surface overlapping model catalogs, but the quote tiles use the baseline row from the Z.AI developer pricing page, not a reseller or proxy price.
Q · 05 Is there a batch discount? +
No separate batch discount row is listed on the Z.AI pricing page for GLM-5 Turbo. The quote tiles show real-time list pricing; batch economics should be treated as a separate calculator variant only when the vendor documents it.
Q · 06 How accurate is the tokenizer estimate? +
The browser widget uses a zhipu-tokenizer-estimate chars-per-token estimate for English text. It is useful for rough planning, but actual billing comes from the vendor API usage fields and can differ for Chinese, code, or mixed-language prompts.