Last verified
PREVIOUS GENERATION1M CONTEXTMULTIMODAL + PDF65K MAX OUTPUTHALF PRICE UNTIL 2027

Gemini 3.7 Flash API Pricing

Gemini 3.7 Flash shipped on August 14, 2026 for complex coding, agentic workflows and reliable multi-step execution. Google relabelled it "our previous-generation Flash model" on September 2, 2026, when Gemini 3.8 Flash arrived on the identical rate card — so nothing about the price changed, only the position. It arrives on the same rate card as Gemini 3.6 Flash, which Google moved onto the same schedule the same day: $0.75/M input, $3.75/M output and $0.075/M cached input through December 31, 2026, then standard pricing of $1.50 / $7.50 / $0.15 from January 1, 2027. Because the two are priced identically, the choice between them is about capability, not cost.

Input - per 1M tokens
$0.75/M
$1.50 from Jan 1, 2027 -50%
Output - per 1M tokens
$3.75/M
$7.50 from Jan 1, 2027 -50%
Cached input - per 1M tokens
$0.075/M
$0.15 from Jan 1, 2027 -90%
Effective - agentic blend
$0.481/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with the Gemini 3.7 Flash rates in effect today — $0.75 input, $3.75 output, $0.075 cached. All three double on January 1, 2027, so a workload sized here costs twice as much from that date. Gemini 3.6 Flash prices identically, so switching between the two changes capability rather than the bill.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · gemini-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 642
Words 100
Tokens (estimated) 167 tokens
Cost as input · uncached $0.00013 USD
Cost as output · uncached $0.00063 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Gemini 3.7 Flash Current $0.75 cache $0.075 $3.75 $0.481 previous-generation Flash 1M Coding, agents, multi-step execution
Gemini 3.6 Flash $0.75 cache $0.075 $3.75 $0.481 same price, prior gen 1M Previous-generation Flash
Gemini 3.5 Flash $1.50 cache $0.15 $9.00 $1.08 pricier output 1M The model 3.6 replaces
Gemini 3.1 Pro Preview $2.00 cache $0.20 $12.00 $1.44 pricier 1M Google's frontier reasoning tier
Gemini 3.5 Flash-Lite $0.30 cache $0.03 $2.50 $0.272 cheaper 1M High-volume, simpler tasks
Gemini 3.1 Flash-Lite $0.25 cache $0.025 $1.50 $0.18 cheapest Gemini 1M Cheapest per-token Gemini tier
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cheaper 200K Anthropic's fast tier
GPT-5.4 mini $0.75 cache $0.075 $4.50 $0.541 cheaper 400K OpenAI's mid tier
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 cheaper 1M Budget reasoning and coding

Frequently asked.

Gemini 3.7 Flash pricing questions, with the dated schedule kept separate from the rate in effect today.

Q · 01 How much does Gemini 3.7 Flash cost? +
Google lists gemini-3.7-flash at $0.75/M input, $3.75/M output including thinking tokens, and $0.075/M cached input. Those rates run through December 31, 2026; from January 1, 2027 the standard pricing is $1.50 / $7.50 / $0.15.
Q · 02 Is Gemini 3.7 Flash more expensive than 3.6 Flash? +
No — they are priced identically, down to the cached rate and the January step-up. Google shipped 3.7 and moved 3.6 onto the same schedule on the same day. Under our standard blend both land at $0.48/M effective, so the decision between them is a capability decision.
Q · 03 What does the January 1, 2027 change mean for my bill? +
Every rate doubles. A workload costing $1,000 a month at today's rates costs $2,000 from that date with no change on your side. Cache storage moves too, from $0.50 to $1.00 per 1M tokens per hour. We re-verify this page against Google's pricing page and will record what actually happens on the date rather than assuming the schedule holds.
Q · 04 What can Gemini 3.7 Flash take as input? +
Text, images, video, audio and PDF, returning text. The model page lists a 1,048,576-token input limit and a 65,536-token output limit, with context caching supported.
Q · 05 Are thinking tokens billed separately? +
No — Google's output price is explicitly "including thinking tokens", so reasoning is billed at the same $3.75/M as visible output. There is no separate reasoning rate to budget for, but heavy thinking still raises the output token count.