AI Model Price Tracker
Every major model's current API price in one table — list rates, cached-input rates, and what the two actually add up to on a real workload — plus a running changelog of cuts, hikes, and new releases. Last verified: August 2026.
Current prices (per 1M tokens, standard tier)
| Model | Provider | Input | Output | Cached input † | Tier |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.01 | Budget | |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.02 | Budget |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | $0.02 | Budget |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.03 | Budget | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | Fast |
| Gemini 3.6 Flash | $1.50 | $7.50 | $0.15 | Fast | |
| Gemini 3.7 Flash | $1.50* | $7.50* | $0.15 | Fast | |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | Fast | |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | Flagship |
| Gemini 3.1 Pro | $2.00 | $12.00 | $0.20 | Flagship | |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $0.20 | Flagship |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | $0.25 | Flagship |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $0.50 | Frontier |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $0.50 | Frontier |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | $0.50 | Frontier |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $0.50 | Frontier |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | Frontier |
*Gemini 3.7 Flash has an introductory rate of $0.75 / $3.75 through December 31, 2026 — list price shown. Estimate your own workload with the AI cost calculator.
Changelog
August 17, 2026 — Sonnet 5's intro rate made permanent, Gemini 3.7 Flash launches
- Price change: Claude Sonnet 5 (Anthropic): $3.00 / $15.00 → $2.00 / $10.00 — Anthropic cancelled the September 1 reversion to list price and made the $2/$10 introductory rate permanent (announced August 11).
- New: Gemini 3.7 Flash (Google) — launched August 13 at an introductory $0.75 / $3.75 through December 31, 2026; standard rate afterward is $1.50 / $7.50, a price twin of Gemini 3.6 Flash. Gemini 3.6 Flash remains available.
August 3, 2026 — OpenAI price cuts
- Cut: GPT-5.6 Luna (OpenAI): $1.00 / $6.00 → $0.20 / $1.20 — an 80% cut, moving Luna from the Fast tier into Budget and making it the cheapest OpenAI model tracked here.
- Cut: GPT-5.6 Terra (OpenAI): $2.50 / $15.00 → $2.00 / $12.00 — a 20% cut; Terra now matches Gemini 3.1 Pro's price exactly instead of GPT-5.4's.
- OpenAI made both cuts on July 30, 2026, three weeks after the GPT-5.6 family's July 9 launch.
July 27, 2026 — two new model launches
- New: Claude Opus 5 (Anthropic) — $5.00 / $25.00, same price as Opus 4.8 but a generational capability jump; Opus 4.8 remains available.
- New: Gemini 3.6 Flash (Google) — $1.50 / $7.50, an output-price cut from Gemini 3.5 Flash's $9.00 with better token efficiency; 3.5 Flash remains available.
- No changes to prices on any previously tracked model this cycle.
July 2026 — tracker launched (baseline)
- Baseline prices recorded for the 14 models tracked at launch. Two more have been added since — see the entries above.
- Notable going in: OpenAI's new GPT-5.6 family (Luna/Terra/Sol) spans $1–$5 input; Claude Sonnet 5 is running an intro discount ($2/$10) through the end of August; Gemini 2.5 Flash-Lite remains the cheapest tracked model at $0.10 input.
- This page is updated monthly — bookmark it to catch cuts before your next bill.
The one trend to know
The price of a given level of AI capability has been falling fast — industry analyses put it around 10× cheaper per year for equivalent quality, driven by hardware, efficiency gains, and competition. Practically: the flagship model you priced six months ago probably has a successor that's cheaper and better. Re-check before committing an annual budget, and design your product so the model can be swapped easily.
List price vs. what you actually pay
The table above is the ceiling. Two mechanisms move a real bill well below it, and which one applies depends entirely on whether a human is waiting for the answer.
Prompt caching discounts input the model has already seen — a long system prompt, an attached document, or the conversation so far. It suits anything real-time and repetitive. Batch processing halves the price of work submitted as a job and collected later, typically within hours. It suits anything nobody is waiting on. They rarely both apply to the same workload.
| Workload | Model | List price | Optimised | Saving |
|---|---|---|---|---|
| Support chatbot, 1,000 conversations/month real-time — caching applies, batch does not | Claude Haiku 4.5 | $16.00 | ~$7.36 with caching | ~54% |
| Daily meeting-notes summaries, one year nobody waiting — batch applies, caching adds little | GPT-5.4 | $7.25 | ~$3.63 with batch | ~50% |
Chatbot assumption: 1,000 conversations × 8 turns ≈ 12M input and 0.8M output tokens, of which about 80% of input is re-sent history eligible for a cache hit. Notes assumption: 250 workdays × 8,000 input and 600 output tokens. Full workings for these and six other jobs are on the cost benchmarks page.
The practical read: if your workload is conversational, caching is the first thing to implement and usually beats switching to a cheaper model. If your workload is asynchronous, batch is close to free money and takes an afternoon to adopt. Reaching for a smaller model before doing either is the common ordering mistake.
How to read AI pricing pages without getting burned
- Input vs output: output is 3–6× dearer; workloads that write a lot (drafting, code generation) cost more than the input price suggests.
- Long-context surcharges: some models charge more above ~200K tokens of context.
- Cached/batch rates: the headline price is the ceiling — caching (~90% off repeated input) and batch (~50% off) are the working floor. Worked examples are in list price vs. what you actually pay above, and eight full job costings are on the cost benchmarks page.
- Intro pricing: launch discounts expire (see Sonnet 5 above) — note the end date in your cost model.
Frequently asked questions
How often is this updated?
Monthly, and after any major provider announcement. The "last verified" date at the top tells you the snapshot's freshness.
Why do some sites show different prices?
Usually one of: outdated data, batch/cached rates presented as headline prices, reseller markups, or regional/enterprise tiers. We track the providers' standard public list prices.
Do prices differ by region?
List prices are global in USD; cloud-marketplace versions (AWS Bedrock, Google Vertex, Azure) can differ slightly and add enterprise features.
The guide that goes deeper
You might also need
🤖AI API Cost Calculator
Compare GPT, Claude & Gemini pricing and estimate monthly costs.
💠AI Pricing by Provider
Every model priced, compared within each provider lineup.
🔢Token Calculator
Token count and cost of any text, on every model.
💬Chatbot Cost Simulator
Monthly LLM bill for your product: users × messages × tokens.