Is DeepSeek's new V4 pricing really about margin expansion, or are they just running out of compute?
With DeepSeek's new pricing taking effect on August 17, there’s been a lot of frustration about the sharp rate increases—especially V4 Pro jumping to 27 RMB ($3.96) / 1M output tokens at peak, and cache hits increasing significantly. Many comments seem to frame this as standard corporate profit-seeking now that they have market share. But looking at how the new pricing is actually structured, does that narrative really make sense? A few details make me wonder if this is almost entirely about compute capacity and load control: Why use peak/off-peak pricing instead of a flat rate hike? The higher rates only apply to 7 hours of the day (9–12 and 14–18 Beijing time), while the other 17 hours are 50% off. If the goal were simple revenue maximization, wouldn't a flat increase across all hours be far more effective? This looks much more like an attempt to push batch jobs and asynchronous workloads into off-peak hours to keep the servers from melting. C-suite/Consumer side remains 100% free: If a company wants quick monetization, gating web/app access with a small monthly subscription usually converts much faster than heavily taxing API developers. Hardware limits are real: With massive token consumption from Agent-driven workflows and obvious hardware supply constraints, isn't it likely their clusters just hit a hard physical ceiling during workday hours? Is it possible that treating the API as an elastic traffic throttle is simply their only viable way to prevent severe queue degradation right now? Curious how those of you working on infra and high-throughput pipelines see this. Please do not forget that, at present, China is essentially unable to obtain Nvidia's high-end computing cards (AI chips) through official channels. To quote a Chinese joke: What use are users? Don't disturb Uncle Liang while he's training AGI!