◆ AI Article Pricer Compare AI article generation costs
Transparent methodology

How AI article costs are calculated

See every assumption behind the calculator—from words and drafts to provider rates, discounts and the costs we cannot predict.

Pricing inputs checked weekly Last provider-rate verification: 2026-09-24
0220% overhead

on input for instructions and scaffolding

03One API call

for each complete draft generated

04Published rates

applied separately to input and output

The calculation

From article brief to cost estimate

The model follows the same four steps for every provider. Only the published token rates and eligible discounts change.

1

Count words

Article length, brief length and any previous drafts.

2

Convert to tokens

Words × your tokens-per-word setting.

3

Separate usage

Input and output tokens are calculated independently.

4

Apply rates

Token totals × the model's current per-million rates.

Exact formula
output = words × tokensPerWord × drafts
input  = (brief × drafts + previousDrafts) × (1 + overhead)
cost   = input/1M × inputRate + output/1M × outputRate

What that means in practice

The brief is sent on every call because the model has no memory between requests. With iterative revision on, each previous draft is also sent back for editing. Turn it off when every draft is independently generated from the original brief.

Prompt overhead applies to input only. It represents system prompts, formatting rules, a style guide and other instructions that travel with the request.

Diagram of the two billed token streams: the brief becomes 390 input tokens costing $0.00078, the 1,500-word article becomes 1,950 output tokens costing $0.0195, totalling $0.0203 on Claude Sonnet 5.
The same formula as a pipeline. Only the output stream scales with article length, which is why a longer brief costs far less than a longer article.
Worked example

A 1,500-word article, step by step

Claude Sonnet 5, a 250-word brief, one draft, 1.3 tokens per word and 20% input overhead.

Article output1,500 × 1.31,950 tokens
Brief input250 × 1.3 × 1.2390 tokens
Total usageinput + output2,340 tokens
What revisions change: Three iterative drafts of this article rise to 11,700 tokens and $0.0702, because earlier drafts become input to later calls. The relationship is not simply “one draft × three.”
Provider adjustments

Discounts are applied only where they exist

The calculator never assumes a global discount. It checks whether the selected model and provider actually support the requested tier.

Model-specific savings

Batch

Eligible Claude, OpenAI and Gemini token charges are 50% off. Grok 4.3 and Grok 4.20 Batch is 20% off; other tracked Grok models do not support Batch. DeepSeek uses a separate time-based off-peak discount, not Batch.

OpenAI only

Flex

Slower scheduling at the batch rate for eligible OpenAI models. Useful when completion time is not urgent.

Costs 2×

Fast

Premium processing on supported OpenAI models and Claude Opus 5 / Opus 4.8. Speed varies; Fast and Batch are separate choices.

Input only

Prompt caching

Uses a published cache-read rate when available. Article workflows benefit only when a substantial fixed prompt is repeatedly reused and actually hits the cache.

Estimate boundaries

Where the calculator will differ from your invoice

Use the result as a planning estimate, then reconcile it against your provider dashboard once the workflow is running.

The estimate is intentionally conservative in scope—not a promise of final spend. It prices predictable text tokens and excludes usage that cannot be inferred reliably from article length.
Not modeled

Reasoning tokens

Thinking models can generate invisible reasoning output billed at the normal output rate. The amount varies by task and cannot be predicted from word count.

Not modeled

Retries and rejected drafts

Failed calls, refusals, regenerations and unusable drafts still consume billable tokens. Real production pipelines often retry.

Not modeled

Tools, files and web search

Retrieval, search and tool calls can add input, output and provider-specific per-call fees beyond the article itself.

Approximation

Tokenizer variation

1.3 tokens per word fits typical English prose. Code, markup, punctuation-heavy copy and non-Latin scripts can require 1.5–2.0 or more. Compared with recorded API output usage below, every draft used more.

Long context

Standard rates fit normal article workflows

Some providers charge more above very large context thresholds. A normal article and brief sit far below them, so the calculator uses standard rates and footnotes affected models. If your workflow sends entire research libraries or very large files, verify the provider's long-context tier separately.

Out of scope

Every rate here is a hosted API rate

The whole model prices metered tokens on someone else's servers, which is why cost scales with volume and never with hardware. Running an open-weight model yourself inverts that: the marginal draft is close to free and the real question becomes whether a given GPU can hold the model and its KV cache at all. That sizing problem is a different calculation and we do not attempt it — RunMyLLM's sizing method works it through against specific GPUs and Apple Silicon, and is the right place to start if self-hosting is on the table. The figures on this page are the number to compare it against.

Measured, not assumed

What 1.3 tokens per word missed in recorded API usage

4 briefs × 5 models, 2026-08-20. Every draft is published in full.

All 20 drafts used more output tokens than this formula predicts — between 1.30× and 5.34× the quoted figure for a 1,500-word commission. These are recorded token-usage counts from how-to / procedural, comparison / decision, technical explainer, book chapter (fiction), repriced at the site's current list rates — not preserved invoices.
Model Tokens per written word Range across briefs vs the 1,950 quoted
GPT-5.5 1.50 1.41–1.72 1.95×
Gemini 2.5 Flash 2.41 1.72–2.70 3.50×
Claude Sonnet 5 2.76 1.93–3.59 2.02×
Claude Opus 5 2.81 1.91–3.80 2.45×
GPT-5 nano 2.82 2.23–3.37 4.41×

Mean of 4 published drafts per model, one per brief. The formula above quotes 1,500 × 1.3 = 1,950 output tokens. Input tokens are excluded; on an article workflow they are a small share of the bill.

Error one

Tokens per written word

Measured 1.41 to 3.80 against the 1.3 assumed here. Tokenizer differences account for some of it; reasoning tokens on reasoning models can account for much of it. Those tokens are charged but do not appear in the text you receive.

Error two

Words you did not ask for

Every brief asked for 1,500 words. The drafts run from close to that to nearly double it, and the overshoot is charged at the same rate as the words you wanted. An editor is then paid to cut them, so you pay for the same words twice.

Error three

The same model, a different job

Claude Opus 5 used 1.91 tokens per word on one brief and 3.80 on another — a task effect larger than the gap between several of the models here. GPT-5.5 barely moved (1.41–1.72). Predictability is itself a property worth pricing.

What we changed

The default is still 1.3, and it is a floor

The obvious fix — swap 1.3 for a measured per-model figure — does not survive the third column of that table. The spread is per task as much as per model, so a per-model constant would be a more confident version of the same error. The default therefore stands, stated as a floor with its range beside it. If you are budgeting for real, raise tokens per word under Advanced on the calculator to the top of your model's range rather than its mean, and treat every unadjusted estimate on this site as the cheapest the workflow could possibly be.

Important qualification

AI generation cost is not the cost of finished editorial work

AI API estimate

Prices generation tokens for a draft. It excludes fact-checking, source verification, editing, originality review, brand alignment and publishing.

Human writer estimate

Article words × the per-word rate you enter, defaulting to $0.15. This usually represents finished copy, not raw text generation.

The comparison is an order-of-magnitude benchmark, not a business case for replacing editorial labor. The cheaper draft can still be the more expensive workflow if it creates substantial review and revision work.

Compare AI-assisted production with freelance writing to separate the API subtotal from research, editing and publishing work. Publishing includes being readable by the assistants that now answer for search: an llms.txt file tells them which of your articles matter, at no cost per draft.

Put the method to work

Price your actual article workflow

Change article length, brief size, revision count, publishing volume and provider tier. Every result uses the methodology above and the rates we re-check weekly.