DeepSeek raised API prices by up to 1,100% today, effective in 48 hours — August 16 at 16:00 UTC. The company simultaneously launched V4 Pro out of a four-month preview. The new peak/off-peak billing structure hits cache-heavy workloads hardest: V4 Pro cached input tokens jump from $0.0036 to $0.044 per million during peak hours — a 1,113% increase. If you’re running V4 Flash or V4 Pro in production, you have until Saturday to decide what to do about it.
What Changes on August 16
DeepSeek replaces flat-rate billing with time-of-day pricing. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. Off-peak covers the remaining 17 hours. Off-peak rates are exactly half the peak rate. Here’s the full breakdown from DeepSeek’s official pricing page:
deepseek-v4-flash
| Token Type | Old Rate | Off-Peak | Peak | Peak Increase |
|---|---|---|---|---|
| Cache-hit input | $0.0028/1M | $0.007/1M | $0.014/1M | +400% |
| Cache-miss input | $0.14/1M | $0.22/1M | $0.44/1M | +214% |
| Output | $0.28/1M | $0.66/1M | $1.32/1M | +371% |
deepseek-v4-pro
| Token Type | Old Rate | Off-Peak | Peak | Peak Increase |
|---|---|---|---|---|
| Cache-hit input | $0.0036/1M | $0.022/1M | $0.044/1M | +1,113% |
| Cache-miss input | $0.435/1M | $0.66/1M | $1.32/1M | +203% |
| Output | $0.87/1M | $1.98/1M | $3.96/1M | +355% |
Cache-Heavy Workloads Take the Biggest Hit
The 1,100% headline is accurate but needs context: it applies specifically to V4 Pro cached-input tokens during peak hours. If your application sends long system prompts that previously generated cache hits — think customer support bots, document-processing pipelines, or code review agents — that’s where the pain concentrates. Teams running batch jobs “overnight” in North America are also in for a surprise: those hours fall squarely in the peak UTC window.
For most workloads, the realistic increase is 135–371% at peak, not 1,100%. Still significant. But the developers who have been leaning hardest on DeepSeek’s cache discounts — building their entire cost model around $0.0036/M cached inputs — are looking at a real budget shock. The 17 daily off-peak hours are a real savings lever if you can shift scheduling around them.
“Even after a 2x to 10x price increase, DeepSeek would likely still undercut most Western rivals.”
— Jun Song, DeepSeek founder
He’s not wrong — V4 Pro peak output at $3.96/M is still cheaper than Claude Sonnet 5 at $10/M. But that framing obscures the fact that for high-volume, cache-hit-heavy workloads at peak hours, the effective cost increase is an order of magnitude. Caixin Global reported that DeepSeek processed 8 trillion tokens on a single day in August — 5 trillion of them on free usage. Infrastructure has to be paid for eventually.
The Cheapest Alternatives Right Now
If you’re evaluating alternatives, the cost picture has shifted enough to take a serious look. Intelligent Living’s cost analysis of a 10M token workload shows the new competitive landscape:
- Xiaomi MiMo-V2.5-Pro — $1.94/10M (matches legacy DeepSeek pricing; 10-point intelligence gap vs V4 Pro)
- MiniMax M3 — $2.98/10M blended
- OpenAI GPT-5.6 Luna — ~$3.48/10M with caching ($0.20/$1.20 per million after July 30 cut)
- DeepSeek V4 Pro off-peak — $4.39/10M (still competitive if you can schedule around peak)
- DeepSeek V4 Pro peak — $8.78/10M
- Claude 4.5 Haiku — $10.48/10M
Panic-migrating is probably not necessary. But having a routing fallback is now essential risk management. Eden AI’s breakdown frames it well: “The initial race to zero was unsustainable.” DeepSeek’s pricing was always closer to developer acquisition cost than a sustainable business model. That era is now over.
Three Actions Before August 16
You have roughly 48 hours. Here’s what actually moves the needle:
- Audit your token breakdown. What percentage of your calls are cache hits vs cache misses? If cache hits are under 30%, your effective increase is closer to 200% than 1,100%.
- Shift batch jobs to off-peak. Anything outside 01:00–04:00 UTC and 06:00–10:00 UTC pays half the price. That is 17 hours of daily breathing room.
- Add a routing fallback. Test MiMo-V2.5-Pro or GPT-5.6 Luna against your existing DeepSeek calls. Even a 20% traffic shift to a cheaper model reduces exposure significantly.
- Consider self-hosting V4 Flash. It runs on a single H100. If you have the infrastructure, the economics have changed enough to revisit this option.
Key Takeaways
- DeepSeek API price increase takes effect August 16, 2026 at 16:00 UTC
- The 1,113% figure applies specifically to V4 Pro cached-input tokens at peak hours; typical output token increases are 214–371%
- Peak hours: 01:00–04:00 UTC and 06:00–10:00 UTC; off-peak is 17 hours daily at half price
- DeepSeek remains competitive for most workloads even at new pricing, but the cost advantage over Western rivals has narrowed significantly
- Audit your token mix, shift scheduling where possible, and add a routing fallback before Saturday













