AI & Developer Tools

AI Model Price Tracker

Every major model's current API price in one table — list rates, cached-input rates, and what the two actually add up to on a real workload — plus a running changelog of cuts, hikes, and new releases. Last verified: August 2026.

Current prices (per 1M tokens, standard tier)

ModelProviderInputOutputCached input †Tier
Gemini 2.5 Flash-LiteGoogle$0.10$0.40$0.01Budget
GPT-5.6 LunaOpenAI$0.20$1.20$0.02Budget
GPT-5.4 NanoOpenAI$0.20$1.25$0.02Budget
Gemini 2.5 FlashGoogle$0.30$2.50$0.03Budget
Claude Haiku 4.5Anthropic$1.00$5.00$0.10Fast
Gemini 3.6 FlashGoogle$1.50$7.50$0.15Fast
Gemini 3.7 FlashGoogle$1.50*$7.50*$0.15Fast
Gemini 3.5 FlashGoogle$1.50$9.00$0.15Fast
Claude Sonnet 5Anthropic$2.00$10.00$0.20Flagship
Gemini 3.1 ProGoogle$2.00$12.00$0.20Flagship
GPT-5.6 TerraOpenAI$2.00$12.00$0.20Flagship
GPT-5.4OpenAI$2.50$15.00$0.25Flagship
Claude Opus 4.8Anthropic$5.00$25.00$0.50Frontier
Claude Opus 5Anthropic$5.00$25.00$0.50Frontier
GPT-5.5OpenAI$5.00$30.00$0.50Frontier
GPT-5.6 SolOpenAI$5.00$30.00$0.50Frontier
Claude Fable 5Anthropic$10.00$50.00$1.00Frontier

*Gemini 3.7 Flash has an introductory rate of $0.75 / $3.75 through December 31, 2026 — list price shown. Estimate your own workload with the AI cost calculator.

Changelog

August 17, 2026 — Sonnet 5's intro rate made permanent, Gemini 3.7 Flash launches

August 3, 2026 — OpenAI price cuts

July 27, 2026 — two new model launches

July 2026 — tracker launched (baseline)

The one trend to know

The price of a given level of AI capability has been falling fast — industry analyses put it around 10× cheaper per year for equivalent quality, driven by hardware, efficiency gains, and competition. Practically: the flagship model you priced six months ago probably has a successor that's cheaper and better. Re-check before committing an annual budget, and design your product so the model can be swapped easily.

List price vs. what you actually pay

The table above is the ceiling. Two mechanisms move a real bill well below it, and which one applies depends entirely on whether a human is waiting for the answer.

Prompt caching discounts input the model has already seen — a long system prompt, an attached document, or the conversation so far. It suits anything real-time and repetitive. Batch processing halves the price of work submitted as a job and collected later, typically within hours. It suits anything nobody is waiting on. They rarely both apply to the same workload.

WorkloadModelList priceOptimisedSaving
Support chatbot, 1,000 conversations/month
real-time — caching applies, batch does not
Claude Haiku 4.5$16.00~$7.36
with caching
~54%
Daily meeting-notes summaries, one year
nobody waiting — batch applies, caching adds little
GPT-5.4$7.25~$3.63
with batch
~50%

Chatbot assumption: 1,000 conversations × 8 turns ≈ 12M input and 0.8M output tokens, of which about 80% of input is re-sent history eligible for a cache hit. Notes assumption: 250 workdays × 8,000 input and 600 output tokens. Full workings for these and six other jobs are on the cost benchmarks page.

The practical read: if your workload is conversational, caching is the first thing to implement and usually beats switching to a cheaper model. If your workload is asynchronous, batch is close to free money and takes an afternoon to adopt. Reaching for a smaller model before doing either is the common ordering mistake.

How to read AI pricing pages without getting burned

Frequently asked questions

How often is this updated?

Monthly, and after any major provider announcement. The "last verified" date at the top tells you the snapshot's freshness.

Why do some sites show different prices?

Usually one of: outdated data, batch/cached rates presented as headline prices, reseller markups, or regional/enterprise tiers. We track the providers' standard public list prices.

Do prices differ by region?

List prices are global in USD; cloud-marketplace versions (AWS Bedrock, Google Vertex, Azure) can differ slightly and add enterprise features.

The guide that goes deeper

You might also need

Last reviewed: · Who maintains this · How it is checked

Prices are read from each provider's own published pricing page, not from third-party summaries. A check that runs on every build (check-prices.js) fails the deploy if any two pages on this site quote a model differently.