Peak hours revealed DeepSeek's user geography

The past few weeks have seen plenty of discussion in the US about the spread of Chinese open-weight AI models, with worries that American businesses might switch to them en masse.

But DeepSeek new peak-hour pricing scheme suggests otherwise – the lab core user base still appears to be concentrated in Asia.

Towards the end of July, DeepSeek released an updated version of its lightweight model, dubbed V4-Flash-0731. With just 284 billion parameters, it delivers results comparable to Anthropic Opus 4.8, which is widely believed to run on trillions of parameters.

The initial pricing looked aggressive even by Chinese market standards – $0.14 per million input tokens and $0.28 per million output tokens. Last week, the lineup grew with V4-Pro, available through API and chat at $0.435 per million input tokens and $0.87 per million output tokens.

Based on preliminary benchmarks posted on WeChat, V4-Pro beats Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench. Meanwhile, the lab compute capacity clearly isn't keeping pace with demand – DeepSeek trailed only Anthropic in token volume processed in July, and could take the top spot in the coming months.

That resource crunch is exactly why the company introduced separate pricing for peak and off-peak hours:

  • V4-Flash input – from $0.14 to $0.22 off-peak and $0.44 peak per million tokens

  • V4-Flash output – from $0.28 to $0.66 off-peak and $1.32 peak

  • V4-Pro input – from $0.435 to $0.66 off-peak and $1.32 peak

  • V4-Pro output – from $0.87 to $1.98 off-peak and $3.96 peak

The key detail is hidden in the schedule. DeepSeek peak windows run from 09:00–12:00 and 14:00–16:00 Beijing time, meaning demand climbs during Asian working hours rather than American ones. That points to Western enterprise clients not yet making up the bulk of traffic.

Alibaba is strengthening its own position too, having recently rolled out two Qwen 3.8-generation models. The flagship Qwen 3.8-Max, with 2.8 trillion parameters, ranks second in performance only to Moonshot Kimi K3.

Even more interesting is the compact Qwen 3.8 model at 27 billion parameters, which comes close to Opus 4.5 in coding quality while running on a single MacBook. On the WeirdML test, the Qwen 3.8 2.4T A95B (xhigh) variant scores 75.2%, placing second among open models behind Kimi-K3, though it burns through a lot of reasoning tokens and writes very long code.

As Chinese open-weight models close in on the capabilities of closed solutions from OpenAI and Anthropic, their popularity keeps climbing fast. According to Hugging Face, one of the largest repositories of open-weight models, the Qwen family now accounts for 151,448 derivative downloads.

Tags: