DeepSeek V4 Flash's weekly call volume exceeds 70 trillion tokens, ranking first in the world, with four out of the top five domestic models
According to data from the multi-model aggregation platform OpenRouter, from July 27 to August 2, DeepSeek V4 Flash topped the global AI large model usage ranking with a call volume of 7.22 trillion tokens. Data from the open-source platform OpenCode shows that this model alone processed 8 trillion tokens on August 1. Four out of the top five positions are occupied by Chinese large models: Xiaomi's MiMo-V2.5 (5.1 trillion), Tencent's Hunyuan Hy3 (5.01 trillion), and another DeepSeek V4 Flash model (3.45 trillion) ranked second to fourth, while OpenAI GPT-5.6 Luna (2.99 trillion, a month-on-month increase of 738%) ranked fifth.
The total call volume of Chinese large models last week reached 28.13 trillion tokens, surpassing the United States for 14 consecutive weeks, firmly holding the top position globally. The previously highly anticipated Kimi K3 fell out of the top ten this week, ranking 12th. It is noteworthy that free models are capturing an increasing share of traffic, with some models experiencing month-on-month growth exceeding 50%. Analysis indicates that domestic models have shifted from price competition to relying on inference performance and stable service capabilities, but companies like OpenAI and Anthropic continue to iterate, accelerating the internal competition and the pace of product iteration will determine user retention capabilities.







