01
Separate cache-hit and cache-miss traffic when estimating a repeated-prompt workload.
DeepSeek pricing
Compare current DeepSeek model prices in one place and estimate the cost of your expected input and output token mix. Verify cache-hit rules and the latest official rates before production budgeting.
Current priced models
9
Lowest listed input
$0.140
Deepseek Coder
Lowest listed output
$0.280
Deepseek Coder
Latest catalog verification
Sep 21, 2026
Current catalog
Rates are first-party USD list prices per one million tokens. Current models are ordered by catalog relevance, not by price alone. Marketplace-only prices are excluded so they cannot be mistaken for the model creator's API rate.
Swipe the table sideways to see context and availability.
| Model | Input / 1M | Output / 1M | Context | Provider listings |
|---|---|---|---|---|
| DeepSeek V4 Flash Vision Exp Released 2026-08-21 | $0.300 | $1.20 | 1.0M | 4 |
| DeepSeek V4 Pro 0423 Released 2026-04-24 | $1.32 | $3.96 | 1.0M | 14 |
| DeepSeek V4 Flash 0423 Released 2026-04-24 | $0.300 | $1.20 | 1.0M | 16 |
| DeepSeek V3.2 Released 2025-12-01 | $0.280 | $0.400 | 164K | 9 |
| R1 Released 2025-01-20 | $0.550 | $2.19 | 164K | 13 |
| DeepSeek V3 Released 2024-12-26 | $0.280 | $0.420 | 131K | 11 |
| Deepseek Coder | $0.140 | $0.280 | 128K | 3 |
| Deepseek Flash | $0.300 | $1.20 | 1.0M | 1 |
| Deepseek Reasoner | $0.280 | $0.420 | 131K | 1 |
Estimate
Select a provider and model, enter your token counts, and get an instant cost estimate.
Configure
Step 4 of 4 · Tokens
Listed rates
DeepSeek V4 Flash Vision Exp
Estimated run
$0.0009
From 1,000 in · 500 out
Input / 1M
$0.300
Output / 1M
$1.20
Prices are official list rates per million tokens.
Full pricing detailCost guidance
01
Separate cache-hit and cache-miss traffic when estimating a repeated-prompt workload.
02
Reasoning workloads can generate substantially more output, so use observed completion lengths rather than a generic average.
03
Compare native DeepSeek rates with third-party inference listings only when latency, region, and service terms are equivalent.
Questions
The main drivers are input tokens, output tokens, cache behavior, and the model selected. Reasoning workloads may produce longer outputs than simple chat requests.
Not necessarily. Inference providers can set their own rates and service terms. This page focuses on DeepSeek as the model creator; compare provider listings on individual model pages.