DeepSeek reprices its API at 16:00 UTC on Sunday, and output goes from US$0.87 to US$3.96 per million tokens at peak.
Off-peak is half of that, which is still more than double the rate every stack in the region was built against.
The capability arrived first. V4-Pro-0813 shipped on 13 August with a one-million-token context window, an agent harness released under an MIT licence, and a vendor-reported jump on Terminal Bench 2.1 from 72.1 to 87.9. Three days sat between the model and the invoice.
The line that moved most is the one nobody watches.
Cache hits go from US$0.003625 to US$0.044 per million at peak, close to 1,100 percent. An agent loop is cache-heavy by construction, re-reading the same system prompt and the same tool definitions on every turn.
The workload DeepSeek built this release for is the workload that carries the steepest part of the rise.
The workload DeepSeek built this release for is the workload that carries the steepest part of the rise.
Then there is the clock. Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC, which in Singapore is 09:00 to 12:00 and 14:00 to 18:00. The premium window is the working day, and the discount opens at six in the evening.1
OCBC runs more than thirty internal tools on open-weight models, with Gemma on document summaries, Qwen on coding, and DeepSeek on market analysis. AIonOS built its Indonesian tourism and agriculture products on DeepSeek alongside Indosat Ooredoo Hutchison.
None of those were costed against Sunday’s number.
Congestion is the stated reason, and it is the smaller half of the story. DeepSeek raised US$7.4 billion in June and is reported to be raising up to US$8 billion more at a US$74 billion valuation. A company at that stage has to show a gross margin rather than a download count, and Caixin read the move as a break from the price war among Chinese developers.
DeepSeek stays cheap in absolute terms. RAND put Chinese models at roughly a sixth to a quarter of the cost of comparable American systems earlier this year, and Sunday’s rates leave that gap open.
Cheap is a level, and the thing that moved is the direction.
Cheap is a level, and the thing that moved is the direction.
What survives a repricing is a routing layer, written before the next promotional rate rather than after it. Route by task instead of by vendor: the interactive path, which cannot wait, goes to whichever model is cheapest between 09:00 and 18:00 Singapore time, and the batch path moves into the evening where the discount is real.
Price every automation at the peak rate inside the model, never the promotional one. A margin that only holds during a promotion belongs to the vendor.
Break the cache line out as its own budget item, because that is where the 1,100 percent lands, and it will not show up in a per-request average.
Then test on the workflow rather than the leaderboard. Artificial Analysis puts V4-Pro at 53 on its intelligence index against 63 for Claude Opus 5 and 60 for Kimi K3, which is a different ranking from the one in the launch post.
At 16:00 UTC on Sunday the number changes and nothing else does. The stacks with a switch in them will read it as a configuration change, and the rest will read it in October, on the bill.
Footnotes
-
The evening window is a real discount for a batch job and worth nothing to a person waiting on a reply, which is the distinction the peak hours exist to price. ↩