DeepSeek is moving its API to time-of-day pricing. From 16:00 UTC on August 16 the company will bill peak and off-peak rates, with off-peak set at half of peak, according to its own developer documentation.

Peak covers 01:00 to 04:00 and 06:00 to 10:00 UTC. Those windows are the Chinese working day, roughly 9am to noon and 2pm to 6pm in Beijing. Everything outside them is off-peak.

The rates are rising as well as splitting. DeepSeek currently charges $0.435 per million input tokens on a cache miss and $0.87 per million output tokens for V4-Pro, with cached input at $0.003625. V4-Flash runs at $0.14 and $0.28, with cached input at $0.0028. Peak output pricing for V4-Pro has been reported at $3.96 per million tokens, and increases across the models, token types and time bands have been put at anywhere from roughly 50 percent to about 1,100 percent.

DeepSeek told developers on August 6 that a significant increase was coming across its API, without giving figures or a date. That warning landed about a week after it launched V4-Flash, marketed as one of the cheapest capable models available anywhere. The rates and the effective date arrived on August 13. If the reported peak figure holds, V4-Pro output rises about four and a half times from its current price, and a developer able to shift the same work off-peak would pay roughly half the peak rate.

The change arrived alongside V4-Pro itself, a 1.6 trillion-parameter mixture-of-experts model with a one million-token context window and reasoning effort adjustable across low, high and maximum settings. It accepts OpenAI’s Responses API format, which lets developers point existing code at it with fewer changes. DeepSeek reported scores of 87.9 on Terminal-Bench 2.1, 83.3 on CyberGym and 62.7 on DeepSWE.

Time-of-day pricing gives developers a reason to move work that is not time-sensitive, batch evaluations, overnight document processing, bulk generation, into the cheaper window.

Other Chinese labs have run into their own capacity ceilings. Moonshot stopped taking new Kimi K3 subscriptions in July when demand outran its compute, closing the door rather than pricing through it. DeepSeek itself raised outside money for the first time in April at a $10 billion valuation, having run without it until then.

Inference prices have generally moved for other reasons. Google’s Gemini 3.5 Flash launched at three times the price of the model it replaced, an increase argued on capability rather than congestion.

The new rates take effect at 16:00 UTC on August 16.

Sources: DeepSeek API Docs, Tech Startups

–
By the Control Plane Editorial Team