AI & Developer Tools

Google API Pricing: The Gemini Lineup Compared

Google fields 6 models at list prices from $0.10/$0.40 to $2.00/$12.00 per million tokens in and out (August 2026). On a real workload the gap between the ends is 22.1×.

The Gemini lineup

Google prices the bottom of the market more aggressively than either competitor and does not currently field a model at the very top of the range. That shapes what Gemini is for: the lineup is strongest where volume matters and the work is mechanical.

ModelTierInput $/1MOutput $/1MOutput multiplier
Gemini 2.5 Flash-Litebudget0.100.404.00×
Gemini 2.5 Flashbudget0.302.508.33×
Gemini 3.6 Flashfast1.507.505.00×
Gemini 3.7 Flash *fast1.507.505.00×
Gemini 3.5 Flashfast1.509.006.00×
Gemini 3.1 Proflagship2.0012.006.00×

Gemini 3.7 Flash is on an introductory rate right now: 0.75 in and 3.75 out, through 31 December 2026. Every table on this page quotes the list price of 1.50/7.50, because that is the number that survives the promotion and the one to budget an annual commitment against. Until then the chatbot workload actually costs $12.00 rather than $24.00. That is enough to change the answer: at the promotional rate Gemini 3.7 Flash is the cheapest fast-tier model tracked here, ahead of Claude Haiku 4.5 at $16.00 — at list price it is not.

The output multiplier runs from 4.00× to 8.33× across this lineup — a wide band, and the reason which rung is cheapest for you depends on how much you generate rather than on the headline rate alone. (What the multiplier means and how to use it.)

What each step up the ladder actually costs

Priced on one real workload — a support chatbot handling 1,000 conversations a month, about 12M tokens in and 800K out:

The expensive step is Gemini 2.5 Flash → Gemini 3.6 Flash at 4.29× — more than a doubling for a single rung. End to end, Gemini 3.1 Pro costs 22.1× what Gemini 2.5 Flash-Lite does on this workload — $32.08 a month more for the same thousand conversations.

The same six jobs, across the Gemini lineup

Six concrete jobs at Google's rates, cheapest highlighted. (All eight jobs, all 17 models.)

JobGemini 2.5 Flash-LiteGemini 2.5 FlashGemini 3.6 FlashGemini 3.7 FlashGemini 3.5 FlashGemini 3.1 Pro
Summarize a 10,000-word report$0.0015$0.0050$0.0229$0.0229$0.0235$0.0314
Answer one support question (with context)$0.0004$0.0013$0.0056$0.0056$0.0058$0.0078
Draft a 1,000-word blog post$0.0006$0.0034$0.0105$0.0105$0.0125$0.0167
Review a 500-line code file$0.0010$0.0041$0.0165$0.0165$0.0177$0.0236
Extract fields from 50 invoices$0.0070$0.0307$0.1162$0.1162$0.1275$0.1700
Support chatbot, 1,000 conversations/month$1.52$5.60$24.00$24.00$25.20$33.60

Read across a row rather than down a column. On this lineup the choice of model matters most for generation-heavy work: drafting a blog post spans 29.6× from cheapest to dearest, against 21.1× for summarising a report. If your product mostly writes, the rung you pick is the biggest lever you have.

On Gemini 2.5 Flash-Lite an 80% cache hit takes the chatbot workload from $1.52 to about $0.6560; on Gemini 3.1 Pro, from $33.60 to about $16.32 — which is below what Gemini 3.5 Flash costs uncached. In other words, caching on the top rung of this lineup beats dropping to the rung below it without caching.

Price twins inside the lineup

Gemini 3.6 Flash and Gemini 3.7 Flash are priced identically, at $1.50 in and $7.50 out. Every cost figure on this page applies to both members of that pair without adjustment, which means there is no cost argument to be had between them. Choose on capability, latency, context window or rate limits instead. Cost re-enters the decision only if one of them caches or batches better for your particular traffic shape — worth measuring rather than assuming.

Where Google sits against the other providers

Within each tier Google competes with the same workload priced on OpenAI and Anthropic. On the chatbot job:

TierBest Google optionCheapest anywhereGap
budgetGemini 2.5 Flash-Lite — $1.52cheapest in tier
fastGemini 3.6 Flash — $24.00Claude Haiku 4.5 (Anthropic) — $16.001.50× more
flagshipGemini 3.1 Pro — $33.60Claude Sonnet 5 (Anthropic) — $32.001.05× more

Gemini 2.5 Flash-Lite is the cheapest model tracked on this site outright, which is Google's clearest structural advantage: at the volume end there is nothing undercutting it.

Standard-tier list prices, short context. Last checked against Google's own published pricing page on . Cached-input rates, the full 17-model table and the price changelog · all eight benchmark jobs.

Where to draw the line in Google's lineup

Google publishes 6 models here, spanning 22.1× from Gemini 2.5 Flash-Lite at the bottom to Gemini 3.1 Pro at the top on the chatbot workload. That spread is what makes routing worth the engineering here; the general rules are on the tracker, and what follows is where Google's own line falls.

The step worth arguing about is Gemini 2.5 Flash → Gemini 3.6 Flash. Every other rung on this ladder is a 3.7×/1.0×/1.1×/1.3× move; that one is 4.29×, or $18.40 a month on the chatbot workload. Test whether your hard requests actually need Gemini 3.6 Flash before making it the default, because that single decision costs more than every other choice in this lineup combined.

One Google model wins both shapes: Gemini 2.5 Flash-Lite is cheapest for generation-heavy work ($0.0006) and for context-heavy work ($0.0015) alike, so there is no workload where a different Google rung is the cheaper answer. Ranking within Google is therefore stable — the only crossings on this site happen between providers, where output multipliers differ, and that is what the tier table above is for.

Estimate your own mix with the AI API cost calculator, or price a specific piece of text with the token calculator.

Frequently asked questions

Which Gemini model is cheapest?

Gemini 2.5 Flash-Lite, at $0.10 per million input tokens and $0.40 per million output. It is cheapest on all six jobs above.

Is Gemini actually the cheapest option?

At the bottom of the range, yes — Gemini 2.5 Flash-Lite is the cheapest model tracked on this site on both input and output. Whether it is cheapest for your workload depends on how much you generate, since output rates rise faster than input rates as you move up any lineup.

What does the Flash / Pro split mean for cost?

Flash is the volume tier and Pro is the capability tier, and the gap between them is the main pricing decision inside the lineup. The job table above shows where that gap actually bites: it is much wider on generation-heavy work than on summarisation.

How current are these prices?

Last checked against Google's own published pricing page — not a third-party summary — on 17 August 2026. A build check fails the deploy if any two pages here quote a model differently.

Last reviewed: · Who maintains this · How it is checked

Prices are read from each provider's own published pricing page, not from third-party summaries. A check that runs on every build (check-prices.js) fails the deploy if any two pages on this site quote a model differently.