GLM-5.3-Flash topped the B.AI model call volume rankings, with a cumulative throughput exceeding 2.41 trillion Tokens
GLM-5.3-Flash has become the most frequently used and popular model on the B.AI platform, with a cumulative token throughput exceeding 2.41 trillion.
As the first native multimodal model in the GLM-5 series, GLM-5.3-Flash features a total of 320 billion parameters and 18 billion active parameters, employing a hybrid architecture that combines sparse and linear attention, supporting 1 million ultra-long contexts, while ensuring rapid response, powerful reasoning, and high cost-effectiveness.
Starting today, developers can still call this model for free through the B.AI platform, covering diverse scenarios such as high-frequency APIs, code writing, complex agents, and ultra-long document processing.
Related tags






