How AI article costs are calculated
See every assumption behind the calculator—from words and drafts to provider rates, discounts and the costs we cannot predict.
per English word by default — a floor, measured
on input for instructions and scaffolding
for each complete draft generated
applied separately to input and output
From article brief to cost estimate
The model follows the same four steps for every provider. Only the published token rates and eligible discounts change.
Count words
Article length, brief length and any previous drafts.
Convert to tokens
Words × your tokens-per-word setting.
Separate usage
Input and output tokens are calculated independently.
Apply rates
Token totals × the model's current per-million rates.
output = words × tokensPerWord × drafts
input = (brief × drafts + previousDrafts) × (1 + overhead)
cost = input/1M × inputRate + output/1M × outputRate
What that means in practice
The brief is sent on every call because the model has no memory between requests. With iterative revision on, each previous draft is also sent back for editing. Turn it off when every draft is independently generated from the original brief.
Prompt overhead applies to input only. It represents system prompts, formatting rules, a style guide and other instructions that travel with the request.
A 1,500-word article, step by step
Claude Sonnet 5, a 250-word brief, one draft, 1.3 tokens per word and 20% input overhead.
1,500 × 1.31,950 tokens250 × 1.3 × 1.2390 tokensinput + output2,340 tokensDiscounts are applied only where they exist
The calculator never assumes a global discount. It checks whether the selected model and provider actually support the requested tier.
Batch
Eligible Claude, OpenAI and Gemini token charges are 50% off. Grok 4.3 and Grok 4.20 Batch is 20% off; other tracked Grok models do not support Batch. DeepSeek uses a separate time-based off-peak discount, not Batch.
Flex
Slower scheduling at the batch rate for eligible OpenAI models. Useful when completion time is not urgent.
Fast
Premium processing on supported OpenAI models and Claude Opus 5 / Opus 4.8. Speed varies; Fast and Batch are separate choices.
Prompt caching
Uses a published cache-read rate when available. Article workflows benefit only when a substantial fixed prompt is repeatedly reused and actually hits the cache.
Where the calculator will differ from your invoice
Use the result as a planning estimate, then reconcile it against your provider dashboard once the workflow is running.
Reasoning tokens
Thinking models can generate invisible reasoning output billed at the normal output rate. The amount varies by task and cannot be predicted from word count.
Retries and rejected drafts
Failed calls, refusals, regenerations and unusable drafts still consume billable tokens. Real production pipelines often retry.
Tools, files and web search
Retrieval, search and tool calls can add input, output and provider-specific per-call fees beyond the article itself.
Tokenizer variation
1.3 tokens per word fits typical English prose. Code, markup, punctuation-heavy copy and non-Latin scripts can require 1.5–2.0 or more. Compared with recorded API output usage below, every draft used more.
Standard rates fit normal article workflows
Some providers charge more above very large context thresholds. A normal article and brief sit far below them, so the calculator uses standard rates and footnotes affected models. If your workflow sends entire research libraries or very large files, verify the provider's long-context tier separately.
Every rate here is a hosted API rate
The whole model prices metered tokens on someone else's servers, which is why cost scales with volume and never with hardware. Running an open-weight model yourself inverts that: the marginal draft is close to free and the real question becomes whether a given GPU can hold the model and its KV cache at all. That sizing problem is a different calculation and we do not attempt it — RunMyLLM's sizing method works it through against specific GPUs and Apple Silicon, and is the right place to start if self-hosting is on the table. The figures on this page are the number to compare it against.
What 1.3 tokens per word missed in recorded API usage
4 briefs × 5 models, 2026-08-20. Every draft is published in full.
| Model | Tokens per written word | Range across briefs | vs the 1,950 quoted |
|---|---|---|---|
| GPT-5.5 | 1.50 | 1.41–1.72 | 1.95× |
| Gemini 2.5 Flash | 2.41 | 1.72–2.70 | 3.50× |
| Claude Sonnet 5 | 2.76 | 1.93–3.59 | 2.02× |
| Claude Opus 5 | 2.81 | 1.91–3.80 | 2.45× |
| GPT-5 nano | 2.82 | 2.23–3.37 | 4.41× |
Mean of 4 published drafts per model, one per brief. The formula above quotes 1,500 × 1.3 = 1,950 output tokens. Input tokens are excluded; on an article workflow they are a small share of the bill.
Tokens per written word
Measured 1.41 to 3.80 against the 1.3 assumed here. Tokenizer differences account for some of it; reasoning tokens on reasoning models can account for much of it. Those tokens are charged but do not appear in the text you receive.
Words you did not ask for
Every brief asked for 1,500 words. The drafts run from close to that to nearly double it, and the overshoot is charged at the same rate as the words you wanted. An editor is then paid to cut them, so you pay for the same words twice.
The same model, a different job
Claude Opus 5 used 1.91 tokens per word on one brief and 3.80 on another — a task effect larger than the gap between several of the models here. GPT-5.5 barely moved (1.41–1.72). Predictability is itself a property worth pricing.
The default is still 1.3, and it is a floor
The obvious fix — swap 1.3 for a measured per-model figure — does not survive the third column of that table. The spread is per task as much as per model, so a per-model constant would be a more confident version of the same error. The default therefore stands, stated as a floor with its range beside it. If you are budgeting for real, raise tokens per word under Advanced on the calculator to the top of your model's range rather than its mean, and treat every unadjusted estimate on this site as the cheapest the workflow could possibly be.
AI generation cost is not the cost of finished editorial work
Prices generation tokens for a draft. It excludes fact-checking, source verification, editing, originality review, brand alignment and publishing.
Article words × the per-word rate you enter, defaulting to $0.15. This usually represents finished copy, not raw text generation.
The comparison is an order-of-magnitude benchmark, not a business case for replacing editorial labor. The cheaper draft can still be the more expensive workflow if it creates substantial review and revision work.
Compare AI-assisted production with freelance writing to separate the API subtotal from research, editing and publishing work. Publishing includes being readable by the assistants that now answer for search: an llms.txt file tells them which of your articles matter, at no cost per draft.
Price your actual article workflow
Change article length, brief size, revision count, publishing volume and provider tier. Every result uses the methodology above and the rates we re-check weekly.