AI & Developer Tools

AI Cost Benchmarks: What Real Tasks Actually Cost

Published prices are quoted per million tokens, which tells you almost nothing about your bill. This page converts those prices into the thing you actually want to know: what does one real job cost? Eight common tasks, priced on all 17 major models.

Calculated from published list prices. Last verified: August 2026. Every token assumption is shown below — check our arithmetic.

Cost per task, by price tier

One model from each tier, so you can see the shape of the trade-off before reading the full matrix. Figures are for a single run of each task, except the last two, which are already at scale.

Task
tokens in / out
Gemini 2.5 Flash-Lite
Budget
Claude Haiku 4.5
Fast
GPT-5.4
Flagship
Claude Opus 4.8
Frontier
Summarize a 10,000-word report
13,300 in / 400 out
$0.0015$0.015$0.039$0.076
Answer one support question, with knowledge-base context
3,000 in / 150 out
$0.0004$0.0037$0.0097$0.019
Draft a 1,000-word blog post
250 in / 1,350 out
$0.0006$0.0070$0.021$0.035
Review a 500-line code file
7,000 in / 800 out
$0.0010$0.011$0.030$0.055
Translate a 2,000-word article
2,700 in / 3,000 out
$0.0015$0.018$0.052$0.088
Extract structured fields from 50 invoices
40,000 in / 7,500 out
$0.0070$0.077$0.213$0.388
Run a support chatbot — 1,000 conversations a month
12,000,000 in / 800,000 out
$1.52$16.00$42.00$80.00
Summarize daily meeting notes for a year
2,000,000 in / 150,000 out
$0.260$2.75$7.25$13.75

What these numbers actually tell you

The spread is enormous — and it is not subtle. Running that support chatbot costs $1.52 a month on Gemini 2.5 Flash-Lite, $80.00 on Claude Opus 4.8 in the table above, and $160 on Claude Fable 5, the dearest model we track. End to end that is roughly 105× — the same feature, the same users, a different line on your P&L.

Output is where the money goes. Every model charges several times more for output than input. Look at the two extremes: summarising a 10,000-word report sends 13,300 tokens in and gets 400 back, while drafting a blog post sends only 250 in and produces 1,350. The drafting job moves far less text, yet on Claude Opus 4.8 it costs $0.035 against $0.076 for the summary. Reading is cheap; writing is not.

Multi-turn conversations compound. The chatbot row looks disproportionate because it is: most APIs are stateless, so the entire conversation so far is re-sent with every turn. Eight turns does not cost eight times one turn — it costs closer to thirty, because turn eight carries the whole history. Teams budgeting from a single-message test are usually out by an order of magnitude.

The cheapest model is often enough. Extraction, classification, routing and summarising are largely solved at the budget tier. Frontier models earn their price on hard reasoning, long-horizon agent work, and anything where a subtle mistake is expensive. Paying frontier rates to reformat invoices is the most common way to overspend on AI.

All 17 models, four everyday jobs

Same arithmetic, every tracked model. The first three columns are per run; the chatbot column is a monthly total.

ModelProviderSummarize a 10k-word reportOne support answerDraft a 1,000-word postChatbot, 1,000 convos/mo
Gemini 2.5 Flash-LiteGoogle$0.0015$0.0004$0.0006$1.52
GPT-5.6 LunaOpenAI$0.0031$0.0008$0.0017$3.36
GPT-5.4 NanoOpenAI$0.0032$0.0008$0.0017$3.40
Gemini 2.5 FlashGoogle$0.0050$0.0013$0.0034$5.60
Claude Haiku 4.5Anthropic$0.015$0.0037$0.0070$16.00
Gemini 3.6 FlashGoogle$0.023$0.0056$0.011$24.00
Gemini 3.7 FlashGoogle$0.023$0.0056$0.011$24.00
Gemini 3.5 FlashGoogle$0.024$0.0058$0.013$25.20
Claude Sonnet 5Anthropic$0.031$0.0075$0.014$32.00
Gemini 3.1 ProGoogle$0.031$0.0078$0.017$33.60
GPT-5.6 TerraOpenAI$0.031$0.0078$0.017$33.60
GPT-5.4OpenAI$0.039$0.0097$0.021$42.00
Claude Opus 4.8Anthropic$0.076$0.019$0.035$80.00
Claude Opus 5Anthropic$0.076$0.019$0.035$80.00
GPT-5.5OpenAI$0.079$0.019$0.042$84.00
GPT-5.6 SolOpenAI$0.079$0.019$0.042$84.00
Claude Fable 5Anthropic$0.153$0.037$0.070$160

Claude Sonnet 5 is shown at list price; an introductory rate of $2.00 / $10.00 runs through 31 August 2026, which lowers every figure in its row by about a third. Need different assumptions? The AI API cost calculator takes your own numbers.

The assumptions, in full

These figures are arithmetic, not measurements: list price × token count. That makes them reproducible, and it makes the token counts the part worth arguing with. Here they are.

TaskInput tokensOutput tokensBasis
Summarize a 10,000-word report13,30040010,000 English words in (≈1.33 tokens per word) plus a short instruction; a ~300-word summary out.
Answer one support question, with knowledge-base context3,000150A ~2,000-word help-centre excerpt plus the customer question in; a ~110-word answer out.
Draft a 1,000-word blog post2501,350A short brief in; 1,000 words out. Output-heavy — this is where output pricing bites.
Review a 500-line code file7,000800~500 lines of source in (code averages ~14 tokens per line); a ~600-word review out.
Translate a 2,000-word article2,7003,0002,000 words in. Translations usually run longer than the source, so output exceeds input here.
Extract structured fields from 50 invoices40,0007,50050 documents × ~800 tokens in; 50 × ~150 tokens of JSON out.
Run a support chatbot — 1,000 conversations a month12,000,000800,0001,000 conversations × 8 turns. The conversation so far is re-sent every turn, so input compounds to ~12,000 tokens per conversation — the single most underestimated cost in AI products.
Summarize daily meeting notes for a year2,000,000150,000250 workdays × ~8,000 tokens of transcript in, ~600 tokens of notes out.

Two conversions do most of the work: English prose runs about 1.33 tokens per word, and source code about 14 tokens per line. Both vary with the tokeniser and the subject matter — expect ±15% on any single job. To count tokens in your own text exactly, use the token calculator.

Three things that move the real bill

The table above is the list-price ceiling. In production, three mechanisms move the actual number:

Current list and effective prices for every model are on the AI price tracker.

How to use this page

Find the row closest to your workload and read across. If the cheapest tier is affordable, prototype there and only move up when you can point to a specific failure the bigger model fixes. If the frontier tier is affordable at your volume, the decision is easy and cost is not your constraint. The uncomfortable middle — where the cheap model nearly works — is where caching, batching and a smaller prompt are worth more than a model upgrade.

Then keep an eye on the date at the top. The price of a given level of AI capability has been falling roughly an order of magnitude a year. A budget built on this table will be pessimistic within months, which is the good direction to be wrong in.

Frequently asked questions

Are these measured costs or estimates?

Estimates, and deliberately transparent ones: every figure is published list price multiplied by a stated token count. No hidden benchmark harness, no vendor-supplied numbers. Substitute your own token counts and the arithmetic still holds.

Why is the chatbot row so much more expensive than the others?

Because conversation history is re-sent on every turn. A single reply is cheap; the eighth reply in a thread carries the whole thread with it. Across 1,000 conversations that compounding, not the reply itself, is what you are paying for.

Do these prices include caching or batch discounts?

No — the tables show standard list pricing, which is the ceiling. Prompt caching can cut repeated input by around 90% and batch processing roughly halves eligible work, so a well-optimised production system typically lands well under these figures.

How often is this updated?

It is regenerated whenever tracked prices change, alongside the AI price tracker. The "last verified" date at the top is the honest freshness marker.

The guide that goes deeper

You might also need

Last reviewed: · Who maintains this · How it is checked

Prices are read from each provider's own published pricing page, not from third-party summaries. A check that runs on every build (check-prices.js) fails the deploy if any two pages on this site quote a model differently.