Base text-token estimate, not an invoice. Excludes reasoning beyond the word estimate, output overshoot, cache writes/storage, tools, regional surcharges and editing. Unsupported tiers use Standard and say so. Check the calculation assumptions and budget for human editing.
Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.66 input, $1.98 output, $0.022 cached.
One 1,500-word article costs $0.00246 on DeepSeek V4.1 Flash, DeepSeek's cheapest tracked model, and $0.00824 on DeepSeek V4 Pro, its most expensive. DeepSeek V4.1 Flash has no Batch tier. Availability and discounts differ by model; check the selector above.
Cheapest first. These are list rates applied to a token model — see
the methodology for exactly where that diverges from an invoice.
The 1,000w column is the per-1,000-words figure people quote: $0.00168 on DeepSeek V4.1 Flash at these settings.
How DeepSeek prices
What is specific to this provider
Providers do not price the same way, and the differences change which model wins.
Rate structure
The shape of the lineup
DeepSeek's lineup is the smallest we track — two models — and the flattest:
output costs three times input on V4 Pro and four times on V4.1 Flash, a narrower
premium than Anthropic, OpenAI or Google typically charge. What makes it genuinely different is the clock. DeepSeek is the only
tracked provider whose published prices change with the time of day.
Peak: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. All other hours, including the entire weekend, are off-peak at half price. This applies to input, output and cache reads.
Reference tables quote peak rates; choose the planned off-peak scenario in the calculator.
Processing tiers
What discounts exist
DeepSeek publishes no batch tier, no fast tier and no flex tier.
The off-peak discount is the nearest thing it has to batch — the same 50% off — but it
runs on the provider's clock rather than yours: a pipeline scheduled into the discounted
windows gets it, work that cannot wait for them does not.
Context pricing
Where the rate changes
Both models carry a 1M-token context window and publish one rate across it —
no long-context surcharge of the kind Google and xAI apply. Cache reads are the deepest
rates at roughly 3% of input for these models, against 10% for the tracked Claude models, though an
article workflow is too output-heavy for that to move the bill much.
Watch out
What to check before you budget
The halved off-peak rate is real money at volume, but it is not a rate you
control: the windows are fixed UTC hours that may not match when your pipeline runs or
when your editors work. Budget on the peak rate in the table below and treat whatever
the off-peak windows save as upside, not as the plan.
What a draft looks like
No DeepSeek draft is in the sample galleries yet
Five other models wrote the same commissions under the same rules — unedited, blind,
with recorded usage repriced at current rates. Read them to calibrate what a raw
API draft looks like at any price before you budget on one.
Reference
DeepSeek list prices
USD per million tokens, cheapest output first.
Model
Input
Cached in
Output
Context
DeepSeek V4.1 Flash
Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.15 input, $0.60 output, $0.003 cached.
$0.30
$0.006
$1.20
1,000K
DeepSeek V4 Pro
Peak rate. DeepSeek halves every rate outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; weekends are entirely off-peak) — off-peak this model is $0.66 input, $1.98 output, $0.022 cached.
Answers specific to this provider's rate structure.
Which DeepSeek model is cheapest for article writing?
DeepSeek V4.1 Flash is the cheaper of the two tracked DeepSeek models for a full-length draft. Both have a 1M context window and a narrow output premium — four times input on V4.1 Flash, three times on V4 Pro — so the choice between them is about draft quality more than rate shape.
What is DeepSeek off-peak pricing?
Peak: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. All other hours, including the entire weekend, are off-peak at half price. Input, output and cache-read rates are all halved off-peak. Reference tables show peak rates; the calculator can budget a planned off-peak workload.
Does DeepSeek have a batch API discount?
No. DeepSeek publishes no batch, fast or flex tier. The off-peak windows are the practical equivalent of a batch discount — the same 50% off — but on the provider’s schedule rather than yours: only work that runs inside those UTC hours gets the lower rate.
Is DeepSeek cheaper than Gemini Flash for articles?
Not automatically. Google’s Flash-Lite tier is cheaper per article at list rate, and Gemini models also offer a 50% batch tier that DeepSeek lacks. DeepSeek’s case is the off-peak window and its flat 3× output premium. Compare like for like on the full pricing table, at the hour your pipeline actually runs.
How much does DeepSeek prompt caching save?
On a cache hit DeepSeek charges roughly 3% of the input rate — the deepest cache discount we track. It still moves an article bill very little, because a draft is overwhelmingly output: caching pays off when a large fixed system prompt or style guide is reused across many requests, not because you generate many articles.
The bottom line
DeepSeek is one column in a five-column decision
Price your workflow here, then check it against the other providers before committing.
At article volumes the difference between two reasonable models is usually smaller than a
single hour of editing.