Zhipu AI’s GLM-5.3-Flash Achieves Top Global AI Call Volume
- tech360.tv

- 3 hours ago
- 2 min read
Zhipu AI's GLM-5.3-Flash, previously an anonymous entity known as Ox Alpha, has achieved the highest global AI model call volume. This marks 18 consecutive weeks of Chinese models leading worldwide inference volume. The recent surge saw global AI model traffic reach 113 trillion tokens, a significant weekly increase.

Chinese models collectively accounted for 55.16 trillion tokens, showing a 36.26% rise during a recent analysis period. Conversely, United States models recorded 17.07 trillion tokens, increasing by 76.71%. For the first occasion in this period, the top three positions in the global ranking were occupied solely by Chinese models, indicating a notable shift in the market's distribution of inference demand.
Ox Alpha, referred to as "Niu Lai" among Chinese devs, secured the number one rank with 15.7 trillion tokens, up 36%. The model was made available on OpenRouter recently, featuring a context window of approximately 1.05 million tokens. It facilitates input across text, image, and video formats. And it was initially offered for preview without charge.
Independent researcher Ben Davis conducted evaluations on Ox Alpha, assessing its performance across 10 subtasks of the DeepSWE software engineering benchmark. The model achieved an approximate 80% pass rate in these tests. This performance positioned it ahead of several established models, including Claude Fable 5, which recorded 65%, GLM-5.3 and Grok 4.6 both at 62%, and GPT-5.6 Sol with 52%.
Zhipu AI later publicly identified Ox Alpha as its GLM-5.3-Flash, confirming it as the first natively multimodal release within the GLM-5 family of models. The organisation simultaneously open-sourced this model. Its pricing stands at one-tenth that of GLM-5.3, with a limited-time promotional rate reducing it to one-twentieth. So, in its inaugural week, GLM-5.3-Flash debuted on the chart at number six, capturing 6.16 trillion tokens.
DeepSeek-V4-Flash, which had previously held the leading position, descended to second place, registering 12.3 trillion tokens. Xiaomi's MiMo-V2.5 maintained its third-place standing for a second successive week, with 9.14 trillion tokens. Gemini 3.7 Flash entered the rankings for the first time, achieving 3.95 trillion tokens, an increase of 120% from its initial metrics.
The Gemini 3.7 Flash model is currently priced at USD 0.75 per million input tokens and USD 3.75 per million output tokens. This pricing structure remains in effect until the end of Dec. And certain models, namely GLM 5.2 and DeepSeek-V4-Pro-0423, no longer appeared on the chart this week.
The aggregate data, according to a National Business Daily analysis of OpenRouter data for a recent period, demonstrates a sustained pattern of cost-efficient Chinese models attracting increasing developer traffic globally. This trend suggests a reconfiguration in how demand for AI inference is spread across the market. The situation appears more indicative of a broad market redistribution than an isolated contest between individual companies.
Zhipu AI's GLM-5.3-Flash, known previously as Ox Alpha, achieved the top position in global AI model call volume.
Chinese large language models have led weekly inference volume globally for 18 consecutive weeks.
The top three models in the ranking were Chinese for the first time in the analysis period.
GLM-5.3-Flash, a natively multimodal model, was open-sourced and priced significantly lower than its predecessor.
This shift indicates a broader redistribution of AI inference demand rather than a sole competitive race.
Source: Pandaily


