DeepSeek pricing in 2026 spans three tiers: a free web chat at chat.deepseek.com, the DeepSeek V4.1 Flash API at $0.15/$0.60 per million tokens off-peak, and DeepSeek V4 Pro at $0.66/$1.98 per million tokens off-peak, with both API rates doubling at peak. Both API models ship with a 1 million token context window at no extra charge. Off-peak, that makes V4.1 Flash 13x cheaper on input and 17x cheaper on output than GPT-6.1 Sol, and 27x and 33x cheaper than Claude Opus 5.5. It is no longer the cheapest API, though: since 22 September 2026, OpenAI’s GPT-6 Luna at $0.10/$0.50 undercuts it on both input and output at every hour.
This guide covers every part of DeepSeek pricing in 2026: the V4.1 Flash and V4 Pro API rates, the legacy V3 chat and R1 reasoner costs, the free tier and the peak and off-peak pricing that has applied since 16 August 2026. We also cover OpenRouter and AWS Bedrock pricing, plus how DeepSeek stacks up against ChatGPT, Claude and Gemini. If you just want a single subscription that bundles DeepSeek with the other top models, we cover that too at the end. Our full DeepSeek vs Claude comparison puts the two side by side.
The Key Takeaways
- DeepSeek V4.1 Flash costs $0.15/M input off-peak and $0.30/M at peak (cache miss), with output at $0.60 and $1.20, in a 1M token context window.
- Prices rose on 16 August 2026. V4 Pro now costs $0.66/M input off-peak and $1.32/M at peak, with output at $1.98 and $3.96. The permanent 75% discount is gone.
- DeepSeek’s web chat is free for individual users with no Plus or Pro plan.
- The DeepSeek API has no published free allowance. DeepSeek mentions a granted balance but gives no amount or expiry.
- Cache hits cost a fraction of cache misses, so prompt caching can cut bills by 80% or more.
- DeepSeek is no longer the cheapest API. OpenAI’s GPT-6 Luna costs $0.10/$0.50, less than V4.1 Flash on both input and output even off-peak.
What is DeepSeek pricing?
DeepSeek pricing is pay-per-token: you are charged per million tokens of text the model reads (input) and writes (output). There are no monthly subscriptions on the API and no per-seat fees. Since 16:00 UTC on 16 August 2026 the rate also depends on when you call it. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and every other hour is off-peak at exactly half the peak rate. Current rates start at $0.15 per million input tokens off-peak for DeepSeek V4.1 Flash.
DeepSeek built its name on undercutting the Western labs. Its closest Chinese rival on price is Alibaba’s Qwen, broken down in our Qwen pricing guide. As we covered in our DeepSeek V4 launch breakdown, the company’s stated goal is rock-bottom prices, backed by Huawei Ascend and Cambricon chips. That keeps DeepSeek pricing far below the flagship OpenAI and Anthropic tiers, but not below all of them: OpenAI’s GPT-6 Luna, launched on 22 September 2026, now costs less.
DeepSeek pricing 2026 (free chat, V4 Flash and V4 Pro at peak and off-peak rates)
DeepSeek pricing in 2026 covers a free web chat, two API models (V4.1 Flash and V4 Pro), and a granted balance whose size DeepSeek does not publish. All figures below come from the official DeepSeek pricing page and are in USD per 1,000,000 tokens. Our guide to whether DeepSeek is free covers why there is no paid consumer tier at all.
| Plan | Price | Models | Context / limits | Best for |
|---|---|---|---|---|
| Free web chat | $0 | V4.1 Flash (non-thinking + thinking) | Fair-use throttling during peak hours | Individuals chatting at chat.deepseek.com or in the mobile app |
| V4.1 Flash API | $0.15 / $0.30 in / $0.60 / $1.20 out per M tokens, off-peak / peak | deepseek-flash (V4.1-Flash build) | 1M context, 384K max output | Cheap general-purpose API workhorse |
| V4 Pro API | $0.66 / $1.32 in / $1.98 / $3.96 out per M tokens, off-peak / peak | deepseek-v4-pro | 1M context, 384K max output | Frontier reasoning at a fraction of rivals’ cost |
| Granted balance | $0, amount not published | Both API models | Spent before your topped-up balance | Testing the API before you top up |
DeepSeek web chat at $0 (free for individuals)
The DeepSeek web chat at chat.deepseek.com and the official mobile app are completely free for individual users. There is no Plus plan, no Pro tier, and no paywall on file uploads or long conversations. The only catch is fair-use throttling, so during peak hours you may see “Server Busy” warnings. Free chat sessions run on V4 Flash by default and let you toggle DeepThink to switch into the V4 Flash thinking mode.
DeepSeek V4.1 Flash pricing
DeepSeek V4.1 Flash is the default low-cost API model. It costs $0.15 per million input tokens off-peak on a cache miss and $0.30 at peak, with output at $0.60 and $1.20. Cache hits fall to $0.003 and $0.006 per million. It supports both a non-thinking and a thinking mode under the deepseek-flash ID and ships with a 1 million token context window plus up to 384K tokens of output. This is the model to default to for everything that does not require deep reasoning.
The model ID moved twice in 2026. From 31 July 2026 the deepseek-v4-flash ID served DeepSeek-V4-Flash-0731, the official build that replaced the April preview, and Artificial Analysis lifted V4 Flash from 40 to 50 on the Intelligence Index version it was running at the time. On 10 September 2026 that ID was retired outright: it still resolves, but DeepSeek now serves those requests with DeepSeek-V4.1-Flash under the deepseek-flash name and bills them at the Flash rate. The price moved with it. The flat $0.14 / $0.28 rate this model launched on was retired on 16 August 2026, and the rate card that replaced it was itself cut on 10 September 2026, so any cost-per-task figure calculated before that date no longer holds.
DeepSeek V4 Pro pricing
DeepSeek V4 Pro is the flagship reasoning model. It costs $0.66 per million input tokens off-peak on a cache miss and $1.32 at peak, with cache hits at $0.022 and $0.044, and output at $1.98 and $3.96. V4 Pro shares the same 1M context and 384K max output as V4 Flash, and the model ID is deepseek-v4-pro. It is still far cheaper than any Western flagship, but the 75% discount DeepSeek once called permanent ended in August 2026, and on the Artificial Analysis Intelligence Index v4.3.2 it now scores 36.0, below V4.1 Flash’s 39.5.
DeepSeek granted balance (amount not published)
DeepSeek’s billing rules mention a granted balance and say it is spent before any money you top up. That is all the pricing page says: it gives no amount, no expiry and no eligibility rule, so the figure of 5 million free tokens for 30 days that circulates online has no DeepSeek source behind it. Treat any credit in your console as a bonus and budget on the pay-per-token rates above.
DeepSeek per-million-token detail (cache miss vs cache hit)
| Model | Input (cache miss), off-peak / peak | Input (cache hit), off-peak / peak | Output, off-peak / peak | Context | Notes |
|---|---|---|---|---|---|
| deepseek-flash | $0.15 / $0.30 | $0.003 / $0.006 | $0.60 / $1.20 | 1M | Default low-cost model, serving the V4.1-Flash build |
| deepseek-v4-pro | $0.66 / $1.32 | $0.022 / $0.044 | $1.98 / $3.96 | 1M | Flagship reasoning model |
V3 era launch pricing (deepseek-chat) | $0.27 | $0.07 | $1.10 | 64K | Historic. ID retired 24 July 2026 |
R1 era launch pricing (deepseek-reasoner) | $0.55 | $0.14 | $2.19 | 64K | Historic. ID retired 24 July 2026 |
The pre-V4 deepseek-chat and deepseek-reasoner model IDs are gone. They were retired after 24 July 2026, 15:59 UTC, and DeepSeek’s docs now list exactly two model IDs, deepseek-flash and deepseek-v4-pro. If you still have those strings in a config file, the calls will fail. The V3 and R1 rates below are kept for historical reference only.
V4 Flash vs V4 Pro pricing
DeepSeek V4 Flash is the cheap, fast workhorse. V4 Pro is the deep-thinking model with a much larger active parameter count. V4 Pro costs roughly 4 times more on input and a little over 3 times more on output than Flash, and since V4.1 Flash arrived on 10 September 2026 it no longer buys a higher score: on the Artificial Analysis Intelligence Index v4.3.2, V4.1 Flash scores 39.5 and V4 Pro 36.0. Default to Flash, and test Pro on your own workload before paying 4x more for it.
How does DeepSeek pricing work?
DeepSeek pricing has three layers and you save money on each one.
- Tokens. Every request is metered in input tokens (what you send) and output tokens (what the model generates). One million tokens is roughly 750,000 English words.
- Cache hits. If part of your prompt has been seen recently, DeepSeek serves it from cache at a steep discount (about 98% off the cache-miss rate on V4 Flash). System prompts and shared context benefit the most.
- Peak-hour surcharge, live. The V3 and R1 off-peak discount retired with those endpoints. DeepSeek replaced it on 16 August 2026 with peak and off-peak pricing where peak hours cost 2x the off-peak rate, across all billing items. Peak is defined as 9:00 to 12:00 and 14:00 to 18:00 Beijing time, which is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Budget at the peak rate for anything that runs inside those windows.
You only pay for what you use. There is no minimum spend, no monthly fee, and no charge for context window length itself.
Is DeepSeek pricing free? Free tier and web chat
Yes, DeepSeek has free options. The consumer chat experience at chat.deepseek.com and in the official mobile app is completely free for individual users. There is no Plus plan, no Pro tier, and no paywall on file uploads or long conversations. The only catch is fair-use throttling, so during peak hours you may see “Server Busy” warnings.
For developers there is no published free API allowance. DeepSeek’s billing rules mention a granted balance that is spent before your own top-ups, but give no amount or expiry, so plan on paying the standard pay-per-token rates above from the first call.
DeepSeek pricing on OpenRouter, AWS Bedrock and Azure
If you do not want to use DeepSeek’s own API, several third-party providers host the same models. Pricing varies because each platform adds its own margin, infrastructure cost and routing.
| Provider | Model | Input/M | Output/M | Notes |
|---|---|---|---|---|
| OpenRouter | DeepSeek V4 Pro | from $0.25 | from $1.05 | 24 provider listings, DeepSeek’s own at $0.66 / $1.98 |
| OpenRouter | DeepSeek V4.1 Flash | from $0.04 | from $0.30 | 27 provider listings, DeepSeek’s own at $0.15 / $0.60 |
| OpenRouter | DeepSeek V3.2 | from $0.21 | from $0.31 | Mid-tier legacy option |
| OpenRouter | DeepSeek R1 (reasoning) | $0.70 | $2.50 | Original R1 chain-of-thought model |
| AWS Bedrock | DeepSeek V3.2 | $0.62 | $1.85 | Higher than direct, but enterprise-friendly |
| Azure AI Foundry | DeepSeek (various) | varies | varies | Pricing depends on region and SKU |
OpenRouter rates above were read from its catalogue on 25 September 2026. It no longer lists any free DeepSeek model, and the spread between providers is wide, so check each listing’s price, limits and quantisation before you pick one.
DeepSeek R1 vs V3 pricing
This is one of the most-asked questions about DeepSeek pricing. The short answer: R1 reasons better but costs roughly 5 times more per query. When it lands, we will map out DeepSeek R2 pricing as soon as official rates go public.
| Model | Input/M, off-peak / peak | Output/M, off-peak / peak | Best for |
|---|---|---|---|
DeepSeek V3 (legacy deepseek-chat) | $0.27 | $1.10 | General chat, summaries, writing |
DeepSeek R1 (legacy deepseek-reasoner) | $0.55 | $2.19 | Math, logic, multi-step reasoning |
| DeepSeek V4.1 Flash | $0.15 / $0.30 | $0.60 / $1.20 | Replaces V3, cheaper and bigger context |
| DeepSeek V4 Pro | $0.66 / $1.32 | $1.98 / $3.96 | Replaces R1, frontier reasoning |
R1 generates many more output tokens per request because it produces a chain-of-thought before its final answer. Output is the dearer half of the bill, as our explainer on AI tokens sets out. That widens the price gap in practice. With V4, both Flash and Pro support thinking and non-thinking modes, so you can pick reasoning depth on a per-call basis. For more on the original reasoner, see our full breakdown of DeepSeek-R1 and how it beat OpenAI.
How DeepSeek pricing compares to GPT-6, Claude and Gemini
| Model | Input/M, off-peak / peak | Output/M, off-peak / peak | DeepSeek V4.1 Flash multiplier (off-peak) |
|---|---|---|---|
| DeepSeek V4.1 Flash | $0.15 / $0.30 | $0.60 / $1.20 | 1x |
| DeepSeek V4 Pro | $0.66 / $1.32 | $1.98 / $3.96 | 4.4x / 3.3x |
| GPT-6 Luna | $0.10 | $0.50 | 0.7x / 0.8x |
| GPT-6.1 Sol | $2.00 | $10.00 | 13.3x / 16.7x |
| Claude Sonnet 5.5 | $2.00 | $10.00 | 13.3x / 16.7x |
| Gemini 3.1 Pro | $2.00 (under 200K context) | $12.00 | 13.3x / 20x |
| Claude Opus 5.5 | $4.00 | $20.00 | 26.7x / 33.3x |
| GPT-6 Astra | $10.00 | $50.00 | 66.7x / 83.3x |
Rates are standard short-context API prices from each vendor’s pricing page, read on 25 September 2026. Claude Sonnet 5.5 succeeded Sonnet 5 on 28 September 2026 at the same $2/$10, checked on Anthropic’s pricing page on 29 September. The ranking depends on tier: off-peak, DeepSeek is 13x to 83x cheaper than OpenAI’s and Anthropic’s mid and top tiers, but GPT-6 Luna costs less than V4.1 Flash on both input and output at every hour, and at peak it is 3x cheaper on input.
On benchmarks the two cheap tiers are close. On the Artificial Analysis Intelligence Index v4.3.2, V4.1 Flash scores 39.5 against GPT-6 Luna’s 37.3, so choosing Flash buys a small quality edge at a price 1.2x to 3x higher, depending on the hour and on input versus output. For high-volume workloads like content pipelines, summarization, classification and RAG over big corpora, DeepSeek is still one of the cheapest models that holds up, but price GPT-6 Luna against it before you commit. For the full side-by-side including monthly subscription tiers, see our complete AI pricing comparison. And for a head-to-head on capabilities, accuracy and where each model wins, read our DeepSeek vs ChatGPT comparison.
How to optimize DeepSeek API pricing costs
A few quick wins can cut a DeepSeek bill by 50 to 90 percent.
- Pin your system prompt. Cache hits are 98% cheaper on V4 Flash. Keep the same opening context across requests.
- Pick V4 Flash by default. Only escalate to V4 Pro when the task actually needs reasoning depth.
- Cap your output. You pay per output token, so set
max_tokensto the smallest value that still works. - Keep batch jobs outside Beijing office hours. The expensive windows are 9:00 to 12:00 and 14:00 to 18:00 Beijing time (UTC+8), which is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Scheduling heavy jobs outside them halves the bill, and has done since 16 August 2026.
- Use V4’s 1M context wisely. A long context window does not cost extra in DeepSeek’s pricing model, so you can stuff more retrieval results in instead of paying for repeat round-trips.
- Compare the OpenRouter providers. Third-party hosts sell V4.1 Flash from $0.04 per million input tokens, under DeepSeek’s own off-peak rate, but check each one’s limits and quantisation first.
Want DeepSeek without dealing with API keys? Try Fello AI
Most readers do not actually want to write code against an API. They want to chat with the best model for the task without juggling five subscriptions, five web tabs and five different memory contexts.
Fello AI is the simplest way to do exactly that. One subscription gets you DeepSeek, ChatGPT, Claude, Gemini, Grok, Perplexity and several other top models in a single Mac, iPhone and iPad app for $9.99 per month. There is no per-token billing to track, no usage cap to budget around, and no separate account to manage for each provider. Everything lives in one chat window and one history.
Why Fello AI works well for DeepSeek users
DeepSeek is still one of the cheapest serious models on the API, but the official web chat is throttled and the desktop experience is bare-bones. Fello AI fixes both. You get DeepSeek inside a real Mac-native app with prompt history, file uploads, image analysis, voice input and the ability to switch models mid-conversation when DeepSeek is not the right tool for the next message. If your prompt would be better answered by Claude’s writing or GPT’s reasoning, swap models without losing context.
What does it actually save you
Add up the standard subscription cost of the major models: ChatGPT Plus at $20/month, Claude Pro at $20/month, Google AI Pro at $19.99/month, Perplexity Pro at $20/month and Grok at $30/month. That is roughly $110 per month for individual access. Fello AI compresses the same lineup into $9.99 per month, with DeepSeek thrown in as part of the bundle. With 27,000+ five-star reviews across the App Store and Mac App Store, it is the most-loved AI bundle app in the Apple ecosystem.
If you would rather keep DeepSeek separate from the others, we also have a guide to a dedicated DeepSeek desktop client for Mac.
Conclusion
DeepSeek pricing in 2026 is still among the cheapest in the market, but it is no longer a flat rate, and no longer the cheapest. V4.1 Flash at $0.15/$0.60 off-peak, or $0.30/$1.20 at peak, is 13x cheaper than GPT-6.1 Sol off-peak, but GPT-6 Luna at $0.10/$0.50 undercuts it at every hour. V4 Pro at $0.66/$1.98 off-peak is harder to recommend now that V4.1 Flash outscores it. What changed on 16 August 2026 is that your bill now depends on the clock, and on 10 September 2026 the Flash rate came down again. The web chat stays free, and prompt caching cuts repeat input costs by about 98%.
If you build software, go straight to the official DeepSeek pricing page and start with V4 Flash. If you just want to use DeepSeek alongside ChatGPT, Claude and Gemini in one place, Fello AI is the simplest route.
FAQ
Is DeepSeek free?
Yes for individual users. The web chat at chat.deepseek.com and the official mobile app are free with no Plus or Pro tier. The API is pay-per-token: DeepSeek mentions a granted balance but publishes no free allowance.
How much does the DeepSeek API cost per million tokens?
DeepSeek V4.1 Flash costs $0.15 per million input tokens off-peak and $0.30 at peak on a cache miss, with output at $0.60 and $1.20. V4 Pro is $0.66 to $1.32 in and $1.98 to $3.96 out. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday.
What’s the difference between V4 Flash and V4 Pro pricing?
V4.1 Flash is the cheap workhorse at $0.15/$0.60 per million tokens off-peak. V4 Pro is the flagship reasoning model and costs about 4 times more on input and a little over 3 times more on output, at $0.66/$1.98 off-peak. Both double during peak hours.
Is DeepSeek pricing cheaper than ChatGPT or Claude?
It depends which tier you compare, and when you call. Against GPT-6.1 Sol ($2/$10), V4.1 Flash is 13x cheaper on input and 17x on output off-peak, and against Claude Opus 5.5 ($4/$20) it is 27x and 33x cheaper; at peak those gaps halve. Against OpenAI's cheapest tier the answer has flipped: GPT-6 Luna at $0.10/$0.50 costs less than V4.1 Flash on both input and output at every hour, 1.5x less on input off-peak and 3x less at peak.
What’s DeepSeek pricing on OpenRouter?
OpenRouter does not charge one rate. On 25 September 2026 it listed 27 providers for V4.1 Flash, from $0.04 per million input tokens, with DeepSeek’s own endpoint at $0.15/$0.60, and 24 for V4 Pro. It no longer lists any free DeepSeek model.
Does DeepSeek charge extra for the 1 million token context window?
No. The full 1M context is included in the standard token rate on both V4 Flash and V4 Pro.