AI & DevelopmentTech BusinessNews & Analysis

DeepSeek API Prices Jump Up to 1,100% on August 16

DeepSeek API price increase chart showing peak and off-peak pricing for V4 Flash and V4 Pro

DeepSeek raised API prices by up to 1,100% today, effective in 48 hours — August 16 at 16:00 UTC. The company simultaneously launched V4 Pro out of a four-month preview. The new peak/off-peak billing structure hits cache-heavy workloads hardest: V4 Pro cached input tokens jump from $0.0036 to $0.044 per million during peak hours — a 1,113% increase. If you’re running V4 Flash or V4 Pro in production, you have until Saturday to decide what to do about it.

What Changes on August 16

DeepSeek replaces flat-rate billing with time-of-day pricing. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. Off-peak covers the remaining 17 hours. Off-peak rates are exactly half the peak rate. Here’s the full breakdown from DeepSeek’s official pricing page:

deepseek-v4-flash

Token TypeOld RateOff-PeakPeakPeak Increase
Cache-hit input$0.0028/1M$0.007/1M$0.014/1M+400%
Cache-miss input$0.14/1M$0.22/1M$0.44/1M+214%
Output$0.28/1M$0.66/1M$1.32/1M+371%

deepseek-v4-pro

Token TypeOld RateOff-PeakPeakPeak Increase
Cache-hit input$0.0036/1M$0.022/1M$0.044/1M+1,113%
Cache-miss input$0.435/1M$0.66/1M$1.32/1M+203%
Output$0.87/1M$1.98/1M$3.96/1M+355%

Cache-Heavy Workloads Take the Biggest Hit

The 1,100% headline is accurate but needs context: it applies specifically to V4 Pro cached-input tokens during peak hours. If your application sends long system prompts that previously generated cache hits — think customer support bots, document-processing pipelines, or code review agents — that’s where the pain concentrates. Teams running batch jobs “overnight” in North America are also in for a surprise: those hours fall squarely in the peak UTC window.

For most workloads, the realistic increase is 135–371% at peak, not 1,100%. Still significant. But the developers who have been leaning hardest on DeepSeek’s cache discounts — building their entire cost model around $0.0036/M cached inputs — are looking at a real budget shock. The 17 daily off-peak hours are a real savings lever if you can shift scheduling around them.

“Even after a 2x to 10x price increase, DeepSeek would likely still undercut most Western rivals.”

— Jun Song, DeepSeek founder

He’s not wrong — V4 Pro peak output at $3.96/M is still cheaper than Claude Sonnet 5 at $10/M. But that framing obscures the fact that for high-volume, cache-hit-heavy workloads at peak hours, the effective cost increase is an order of magnitude. Caixin Global reported that DeepSeek processed 8 trillion tokens on a single day in August — 5 trillion of them on free usage. Infrastructure has to be paid for eventually.

The Cheapest Alternatives Right Now

If you’re evaluating alternatives, the cost picture has shifted enough to take a serious look. Intelligent Living’s cost analysis of a 10M token workload shows the new competitive landscape:

  • Xiaomi MiMo-V2.5-Pro — $1.94/10M (matches legacy DeepSeek pricing; 10-point intelligence gap vs V4 Pro)
  • MiniMax M3 — $2.98/10M blended
  • OpenAI GPT-5.6 Luna — ~$3.48/10M with caching ($0.20/$1.20 per million after July 30 cut)
  • DeepSeek V4 Pro off-peak — $4.39/10M (still competitive if you can schedule around peak)
  • DeepSeek V4 Pro peak — $8.78/10M
  • Claude 4.5 Haiku — $10.48/10M

Panic-migrating is probably not necessary. But having a routing fallback is now essential risk management. Eden AI’s breakdown frames it well: “The initial race to zero was unsustainable.” DeepSeek’s pricing was always closer to developer acquisition cost than a sustainable business model. That era is now over.

Three Actions Before August 16

You have roughly 48 hours. Here’s what actually moves the needle:

  • Audit your token breakdown. What percentage of your calls are cache hits vs cache misses? If cache hits are under 30%, your effective increase is closer to 200% than 1,100%.
  • Shift batch jobs to off-peak. Anything outside 01:00–04:00 UTC and 06:00–10:00 UTC pays half the price. That is 17 hours of daily breathing room.
  • Add a routing fallback. Test MiMo-V2.5-Pro or GPT-5.6 Luna against your existing DeepSeek calls. Even a 20% traffic shift to a cheaper model reduces exposure significantly.
  • Consider self-hosting V4 Flash. It runs on a single H100. If you have the infrastructure, the economics have changed enough to revisit this option.

Key Takeaways

  • DeepSeek API price increase takes effect August 16, 2026 at 16:00 UTC
  • The 1,113% figure applies specifically to V4 Pro cached-input tokens at peak hours; typical output token increases are 214–371%
  • Peak hours: 01:00–04:00 UTC and 06:00–10:00 UTC; off-peak is 17 hours daily at half price
  • DeepSeek remains competitive for most workloads even at new pricing, but the cost advantage over Western rivals has narrowed significantly
  • Audit your token mix, shift scheduling where possible, and add a routing fallback before Saturday
ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *