LLM API Pricing: every model, every rate
Across the 55 models tracked here, input runs from $0.05 to $30.00 per million tokens and output from $0.40 to $180.00. Applied to one 1,500-word article that is $0.00080 on GPT-5 nano and $0.3627 on GPT-5.4 Pro.
A lookup table, not a calculator. Rates are transcribed by hand from each provider's own pricing page — no aggregators, no estimates — and every figure below links through to the calculator if you want it applied to your own workflow.
Four columns decide what you actually pay
All rates are USD per million tokens, the unit every provider publishes.
Input
What you send: the brief, instructions, any previous draft. Cheap relative to output, and it does not grow with article length.
Cached in
The discounted input rate on a prompt-cache hit. 4 tracked models publish no cache rate at all, so caching cannot reduce their input cost.
Output
What the model generates. This is the number that decides an article's cost, because a draft is almost entirely output.
Out/in
The output premium. A ratio of 5× means output costs five times input — the higher it is, the more a long article costs relative to a long brief.
Per-million rates hide the thing you are buying
Two models with identical input rates can differ several-fold on a real article, because article generation is overwhelmingly an output workload. The 1,500 words column applies one consistent workflow — one draft, a 250-word brief, 20% prompt overhead — to every model, so the rows are directly comparable. Click any figure to open that model in the calculator.
Anthropic pricing calculator
9 models · tiers: Standard, Batch, Fast · cheapest article $0.0101 on Claude Haiku 4.5
| Model | Input | Cached in | Output | Out/in | Context | 1,500 words |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | $1.00 | $0.100 | $5.00 | 5.0× | 200K | $0.0101 |
| Claude Sonnet 5.5 | $2.00 | $0.200 | $10.00 | 5.0× | 1,000K | $0.0203 |
| Claude Sonnet 5 | $2.00 | $0.200 | $10.00 | 5.0× | 1,000K | $0.0203 |
| Claude Sonnet 4.6 | $3.00 | $0.300 | $15.00 | 5.0× | 1,000K | $0.0304 |
| Claude Opus 5.5 | $4.00 | $0.200 | $20.00 | 5.0× | 1,000K | $0.0406 |
| Claude Opus 5 | $5.00 | $0.500 | $25.00 | 5.0× | 1,000K | $0.0507 |
| Claude Opus 4.8 | $5.00 | $0.500 | $25.00 | 5.0× | 1,000K | $0.0507 |
| Claude Fable 5.1 | $10.00 | $0.250 | $50.00 | 5.0× | 1,000K | $0.1014 |
| Claude Fable 5 | $10.00 | $1.000 | $50.00 | 5.0× | 1,000K | $0.1014 |
OpenAI pricing calculator
26 models · tiers: Standard, Flex, Batch, Fast · cheapest article $0.00080 on GPT-5 nano
| Model | Input | Cached in | Output | Out/in | Context | 1,500 words |
|---|---|---|---|---|---|---|
| GPT-5 nano | $0.05 | $0.005 | $0.40 | 8.0× | 400K | $0.00080 |
| GPT-4.1 nano | $0.10 | $0.025 | $0.40 | 4.0× | 1,048K | $0.00082 |
| GPT-6 Luna Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $0.10 | $0.010 | $0.50 | 5.0× | 1,050K | $0.00101 |
| GPT-4o mini | $0.15 | $0.075 | $0.60 | 4.0× | 128K | $0.00123 |
| GPT-5.6 Luna Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $0.20 | $0.020 | $1.20 | 6.0× | 1,050K | $0.00242 |
| GPT-5.4 nano | $0.20 | $0.020 | $1.25 | 6.3× | 400K | $0.00252 |
| GPT-4.1 mini | $0.40 | $0.100 | $1.60 | 4.0× | 1,048K | $0.00328 |
| GPT-5 mini | $0.25 | $0.025 | $2.00 | 8.0× | 400K | $0.00400 |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | 6.0× | 400K | $0.00907 |
| GPT-4.1 | $2.00 | $0.500 | $8.00 | 4.0× | 1,048K | $0.0164 |
| GPT-5.1 | $1.25 | $0.125 | $10.00 | 8.0× | 400K | $0.0200 |
| GPT-5 | $1.25 | $0.125 | $10.00 | 8.0× | 400K | $0.0200 |
| GPT-6 Sol Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $2.00 | $0.200 | $10.00 | 5.0× | 1,050K | $0.0203 |
| GPT-4o | $2.50 | $1.250 | $10.00 | 4.0× | 128K | $0.0205 |
| GPT-5.6 Terra Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $2.00 | $0.200 | $12.00 | 6.0× | 1,050K | $0.0242 |
| GPT-5.2 | $1.75 | $0.175 | $14.00 | 8.0× | 400K | $0.0280 |
| GPT-5.4 Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $2.50 | $0.250 | $15.00 | 6.0× | 1,050K | $0.0302 |
| GPT-5.6 Sol Promotional rate, published as available at least through 2026-11-21; OpenAI has not announced the rate that follows it. Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $4.00 | $0.400 | $20.00 | 5.0× | 1,050K | $0.0406 |
| GPT-5.5 Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $5.00 | $0.500 | $30.00 | 6.0× | 1,050K | $0.0604 |
| GPT-6 Astra Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session. | $10.00 | $1.000 | $50.00 | 5.0× | 1,050K | $0.1014 |
| GPT-5.6 Cyber Security-focused tier, gated behind the Daybreak program. | $12.50 | $1.250 | $75.00 | 6.0× | 1,050K | $0.1511 |
| GPT-5.5 Cyber Security-focused tier, gated behind the Daybreak program. | $12.50 | $1.250 | $75.00 | 6.0× | 1,050K | $0.1511 |
| GPT-5 Pro | $15.00 | — | $120.00 | 8.0× | 400K | $0.2398 |
| GPT-5.2 Pro | $21.00 | — | $168.00 | 8.0× | 400K | $0.3358 |
| GPT-5.5 Pro | $30.00 | — | $180.00 | 6.0× | 400K | $0.3627 |
| GPT-5.4 Pro | $30.00 | — | $180.00 | 6.0× | 1,050K | $0.3627 |
Google pricing calculator
10 models · tiers: Standard, Flex, Batch · cheapest article $0.00082 on Gemini 2.5 Flash-Lite
| Model | Input | Cached in | Output | Out/in | Context | 1,500 words |
|---|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.010 | $0.40 | 4.0× | 1,000K | $0.00082 |
| Gemini 3.1 Flash-Lite Text, image and video input rate; audio input is $0.50. | $0.25 | $0.025 | $1.50 | 6.0× | 1,000K | $0.00302 |
| Gemini 3.5 Flash-Lite | $0.30 | $0.030 | $2.50 | 8.3× | 1,000K | $0.00499 |
| Gemini 2.5 Flash | $0.30 | $0.030 | $2.50 | 8.3× | 1,000K | $0.00499 |
| Gemini 3.8 Flash Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01. | $0.75 | $0.075 | $3.75 | 5.0× | 1,000K | $0.00760 |
| Gemini 3.7 Flash Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01. | $0.75 | $0.075 | $3.75 | 5.0× | 1,000K | $0.00760 |
| Gemini 3.6 Flash Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01. | $0.75 | $0.075 | $3.75 | 5.0× | 1,000K | $0.00760 |
| Gemini 3.5 Flash | $1.50 | $0.150 | $9.00 | 6.0× | 1,000K | $0.0181 |
| Gemini 2.5 Pro Rate shown is for prompts up to 200K tokens; $2.50 / $15.00 above that. | $1.25 | $0.125 | $10.00 | 8.0× | 1,000K | $0.0200 |
| Gemini 3.1 Pro Rate shown is for prompts up to 200K tokens; $4.00 / $18.00 above that. | $2.00 | $0.200 | $12.00 | 6.0× | 1,000K | $0.0242 |
xAI pricing calculator
8 models · tiers: Standard, Batch · cheapest article $0.00429 on Grok Build 0.1
| Model | Input | Cached in | Output | Out/in | Context | 1,500 words |
|---|---|---|---|---|---|---|
| Grok Build 0.1 Rate shown is for prompts under 200K tokens; $2.00 / $4.00 at or above that. | $1.00 | $0.200 | $2.00 | 2.0× | 256K | $0.00429 |
| Grok 4.3 Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that. | $1.25 | $0.200 | $2.50 | 2.0× | 1,000K | $0.00536 |
| Grok 4.20 Reasoning Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that. | $1.25 | $0.200 | $2.50 | 2.0× | 1,000K | $0.00536 |
| Grok 4.20 Non-Reasoning Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that. | $1.25 | $0.200 | $2.50 | 2.0× | 1,000K | $0.00536 |
| Grok 4.20 Multi-Agent Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that. | $1.25 | $0.200 | $2.50 | 2.0× | 1,000K | $0.00536 |
| Grok 4.7 Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that. | $2.00 | $0.500 | $6.00 | 3.0× | 500K | $0.0125 |
| Grok 4.6 Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that. | $2.00 | $0.500 | $6.00 | 3.0× | 500K | $0.0125 |
| Grok 4.5 Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that. | $2.00 | $0.300 | $6.00 | 3.0× | 500K | $0.0125 |
DeepSeek pricing calculator
2 models · tiers: Standard · cheapest article $0.00246 on DeepSeek V4.1 Flash
| Model | Input | Cached in | Output | Out/in | Context | 1,500 words |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.15 input, $0.60 output, $0.003 cached. | $0.30 | $0.006 | $1.20 | 4.0× | 1,000K | $0.00246 |
| DeepSeek V4 Pro Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.66 input, $1.98 output, $0.022 cached. | $1.32 | $0.044 | $3.96 | 3.0× | 1,000K | $0.00824 |
26 models carry a rate qualification
Promotional and context-dependent rates are footnoted rather than averaged in.
- GPT-6 Astra — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-6 Sol — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-6 Luna — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.6 Sol — Promotional rate, published as available at least through 2026-11-21; OpenAI has not announced the rate that follows it. Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.5 — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.4 — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.6 Terra — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.6 Luna — Prompts over 272K input tokens bill at 2x input / 1.5x output for the whole session.
- GPT-5.6 Cyber — Security-focused tier, gated behind the Daybreak program.
- GPT-5.5 Cyber — Security-focused tier, gated behind the Daybreak program.
- Gemini 3.1 Pro — Rate shown is for prompts up to 200K tokens; $4.00 / $18.00 above that.
- Gemini 2.5 Pro — Rate shown is for prompts up to 200K tokens; $2.50 / $15.00 above that.
- Gemini 3.8 Flash — Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01.
- Gemini 3.7 Flash — Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01.
- Gemini 3.6 Flash — Promotional rate through 2026-12-31; $1.50 / $7.50 from 2027-01-01.
- Gemini 3.1 Flash-Lite — Text, image and video input rate; audio input is $0.50.
- Grok 4.7 — Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that.
- Grok 4.6 — Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that.
- Grok 4.5 — Rate shown is for prompts under 200K tokens; $4.00 / $12.00 at or above that.
- Grok 4.3 — Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that.
- Grok 4.20 Reasoning — Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that.
- Grok 4.20 Non-Reasoning — Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that.
- Grok 4.20 Multi-Agent — Rate shown is for prompts under 200K tokens; $2.50 / $5.00 at or above that.
- Grok Build 0.1 — Rate shown is for prompts under 200K tokens; $2.00 / $4.00 at or above that.
- DeepSeek V4.1 Flash — Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.15 input, $0.60 output, $0.003 cached.
- DeepSeek V4 Pro — Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.66 input, $1.98 output, $0.022 cached.
4 models cannot use prompt caching to cut input cost
GPT-5.5 Pro, GPT-5.4 Pro, GPT-5.2 Pro, GPT-5 Pro. Where a provider publishes no cache-read rate we leave the column empty rather than assuming a percentage of the input rate.
What has changed
Every verification pass is recorded here, whether or not anything moved.
| Date | Change |
|---|---|
| Added Claude Sonnet 5.5, released by Anthropic on 2026-09-28 at Sonnet 5’s rates — $2.00 input / $10.00 output, cache reads $0.20, Batch $1.00 / $5.00, no Fast mode, full 1M context at the standard rate (Anthropic pricing). This was not a full catalog re-verification. 55 models tracked. | |
| Added ten models, transcribed from the Anthropic, OpenAI, Google and xAI pricing pages. Claude Fable 5.1 — $10.00 / $50.00, cache reads $0.25 (0.025× input), Batch at half price, no Fast. GPT-6 Luna — $0.10 / $0.50, cache reads $0.01, Batch and Flex at half price, Fast at 2×, 272K long-context threshold. GPT-5.5 Cyber — $12.50 / $75.00, cache reads $1.25, Standard only, Daybreak-gated. GPT-5.2 Pro — $21.00 / $168.00, Batch at half price. GPT-4.1 nano — $0.10 / $0.40, cache reads $0.025, Batch at half price with no cache discount, Fast at 2×. Gemini 3.8 Flash — promotional $0.75 / $3.75, cache reads $0.075, through 2026-12-31, then $1.50 / $7.50 / $0.15 from 2027-01-01. Gemini 3.1 Flash-Lite — $0.25 / $1.50, cache reads $0.025. Both Gemini models offer Batch and Flex at half price. Grok 4.20 Reasoning, Non-Reasoning and Multi-Agent — $1.25 / $2.50, cache reads $0.20, doubling at 200K prompt tokens, Batch 20% off. This was not a full catalog re-verification. 54 models tracked. | |
| Verification pass. Full re-verification of all 44 tracked models against the five provider pricing pages, read from the raw published tables. No standard rate moved. One service-tier gap corrected: Google now publishes Flex at half the Standard rate on every tracked Gemini model, so Flex is now offered for Gemini here; Google’s 1.8× Priority tier is not modeled. Gemini Batch cache reads were also corrected — Gemini 3.1 Pro, 2.5 Pro, 2.5 Flash and 2.5 Flash-Lite keep the Standard cache-read rate on Batch and Flex, and 3.5 Flash-Lite charges $0.02 — where this calculator had been halving them. 44 models tracked. | |
Added Grok 4.7 — $2.00 input / $6.00 output, cache reads $0.50, doubling to $4.00 / $12.00 at or above 200K prompt tokens; no Batch tier listed (xAI pricing). DeepSeek has retired V4 Flash: the deepseek-v4-flash name is still accepted but served by DeepSeek-V4.1-Flash and billed at its price (DeepSeek pricing). V4 Flash is replaced here by V4.1 Flash at $0.30 input / $1.20 output, cache hits $0.006 at peak, with every rate halved off-peak; old links to V4 Flash resolve to V4.1 Flash, the same way the API does. This was not a full catalog re-verification. 44 models tracked. |
|
| Added Claude Opus 5.5 — $4.00 input / $20.00 output, cache reads $0.20 (0.05× input rather than the usual 0.1×), Batch $2.00 / $10.00, Fast $8.00 / $40.00 — and GPT-6 Sol — $2.00 / $10.00, cache reads $0.20, Batch and Flex at half price, Fast at 2×, with the same 272K long-context threshold as the rest of the 1,050K-window GPT line. Both transcribed from the Anthropic and OpenAI pricing pages. This was not a full catalog re-verification. 43 models tracked. | |
| Targeted service-tier corrections: reconciled model-specific OpenAI Fast, Flex and Batch availability and component rates against its pricing tables; Grok 4.3 Batch is 20% off, while the other tracked Grok model pages mark Batch unsupported; and Claude Opus 5 and Opus 4.8 support Fast at 2× Standard. Corrected DeepSeek weekday/weekend schedule wording and added explicit off-peak budgeting. This was not a full base-rate catalog re-verification. | |
| Verification pass. Full re-verification of all 40 tracked models against the five provider pricing pages: nothing moved. Added GPT-6 Astra, released by OpenAI on 2026-09-03 — $10.00 input / $50.00 output, cache reads $1.00, with the same 272K long-context threshold as the rest of the 1,050K-window GPT line (what an article costs on it). One transcription refinement: DeepSeek’s pricing page specifies that its peak-hour windows apply Monday–Friday, so weekends bill entirely at the half-price off-peak rate; both DeepSeek footnotes now say so. 41 models tracked. | |
| Verification pass. Full re-verification of all 40 tracked models against the five provider pricing pages. Two moved. Anthropic cancelled the Claude Sonnet 5 increase recorded on 2026-08-20: the $2.00 / $10.00 introductory rate is now the standard rate and the 2026-09-01 rise to $3.00 / $15.00 will not happen (what changed). GPT-5.6 Sol fell from $5.00 / $30.00 to $4.00 / $20.00, cache reads $0.50 to $0.40, published as promotional pricing available at least through 2026-11-21 with no successor rate named. Everything else re-verified unchanged. | |
| Added DeepSeek as a fifth provider: V4 Flash and V4 Pro, transcribed from DeepSeek’s published pricing, including the off-peak windows in which every rate is halved. 40 models tracked. | |
| Recorded Anthropic’s announced end of Claude Sonnet 5’s introductory rate — $2.00 / $10.00 per million tokens becomes $3.00 / $15.00 on 2026-09-01 (full breakdown) — and the scheduled end of the promotional Gemini 3.6 Flash and 3.7 Flash rates on 2027-01-01. Both now roll over by date automatically. | |
| Verification pass. Baseline: every tracked rate transcribed and verified against the provider pricing pages of Anthropic, OpenAI, Google and xAI. 38 models. |
These are list prices. Your invoice is a workflow.
Drafts, revision context, prompt overhead, reasoning tokens and processing tier all move the total. The calculator applies your own assumptions to every model on this page. They are also all rental rates — for the other side of that trade, where the cost is hardware rather than tokens, RunMyLLM's catalogue of open-weight models lists what actually fits on a given card and how fast it runs there.
Open the calculator →