Z.AI Sets Year-End Revenue Target of $2.4 Billion, Acknowledges 17-Month Gap to Rival, and Reports 40-Fold Surge in Token Usage After 101% Price Hike

Deep News
12 hours ago

Following the release of its first interim results as a Hong Kong-listed company on August 31, Z.AI (02513.HK) convened an earnings call and a subsequent in-depth analyst session, revealing several key data points beyond the financial statements. For the first time, the company issued official guidance for its Annual Recurring Revenue (ARR), projecting a year-end figure of $2.4 billion, while also benchmarking its progress against a major international competitor.

During the analyst call, the company’s Secretary to the Board, Xiao Lei, provided the first complete breakdown of Z.AI's ARR growth trajectory. He confirmed that as of late August, ARR had reached $1.6 billion, a figure he noted corresponds to the level achieved by Anthropic in March 2025, placing Z.AI roughly 17 months behind. He added that this gap has narrowed from an initial 24 months, marking a seven-month improvement. Xiao Lei also disclosed that Z.AI has only officially published ARR on three occasions, with the first being $250 million at the end of March, the second being $1 billion in July, and the third being today's $1.6 billion figure.

Xiao Lei acknowledged that Z.AI's ARR growth has not been smooth, describing it as pulse-like, primarily due to temporary mismatches in computing power supply. He revealed that at the end of June, ARR was around $530-540 million, a figure previously undisclosed. Regarding the comparison with Anthropic, he stated, "Today's $1.6 billion is equivalent to Anthropic's level in March 2025, so we are still about 17 months behind. However, we have shortened the gap from an initial 24-month deficit by a full seven months." He also clarified that ARR represents management's view of order income, not financial revenue recognition, and noted a ratio between ARR and recognized revenue of roughly 2:1 to 4:1 for Anthropic, depending on the observation point.

Addressing pricing and demand, Xiao Lei reported that while the average API selling price increased by about 101%, token call volume has surged more than 40-fold compared to the start of the year. He further elaborated that for the top ten customers by revenue, average daily usage has grown by an impressive 98-fold year-to-date. This combination of price increase and volume explosion is considered commercially unusual. He highlighted the Coding Plan, a programming subscription product, as an extreme example: after being available for 365 days, its usage grew by 234-fold, despite being in a sales suspension for most of the first half due to computing power constraints. Since reopening sales in July after significant infrastructure improvements, usage grew another 15-fold in August even after a price increase. Xiao Lei attributed these shifts not to deliberate strategy but to the natural evolution driven by enhanced model capabilities.

In response to analyst queries about the open platform's gross margin, which stood at 24.6% in the first half, Xiao Lei provided directional guidance for the next 12 to 18 months. He stated the company’s goal is to strive for a gross margin exceeding 50%. He attributed the currently lower margin to four primary factors: new domestic large-scale computing clusters are still in a ramp-up phase; rapid model iteration inevitably leads to resource redundancy and switching costs; there remains room for further infrastructure optimization, even though unit inference costs have already dropped by roughly 80% in the past six months; and finally, ongoing refinement in the pricing structure for flagship and Flash-tier models to reduce misallocation.

Regarding next-generation model technology, Chief Scientist Tang Jie offered his most detailed explanation to date, unveiling the core direction of "Fully Self-Training." He defined this as models capable of self-directing their entire training process from pre-training to post-training, including self-correction and self-stopping. He connected this concept to the international idea of Recursive Self-Improvement (RSI). Tang Jie also addressed questions on model parameter size, explaining that Z.AI’s choice of a 744 billion parameter base model for the GLM-5 series was determined in January to be used for over half a year. He defended this relative restraint by noting that with domestic training data typically in the 30-50 trillion token range, scaling law suggests that excessively large models would see diminishing returns. He confirmed that the next-generation base model is already in pre-training, and expressed confidence it will be available at full capacity on its launch day.

Sharing user depth data for the first time, Xiao Lei disclosed that the MaaS platform had over 7.4 million registered users by the end of August, a 144% increase year-to-date, and that paying daily active users had grown by 603%. More notably, the top ten users by revenue accounted for nearly 40% of the total daily token usage, exceeding 15 trillion tokens. He also revealed that four of China’s top ten internet companies have designated GLM as their primary global model choice. For the Coding Plan, which has been live for 365 days, usage grew 234-fold, positioning it as one of the world's largest and fastest-growing programming services.

On international expansion, Xiao Lei indicated that the company is actively pursuing collaborations with overseas cloud service providers, expecting significant progress within the next one to two months. He mentioned initial cooperation with AWS on distribution and ongoing discussions with two major Western cloud providers and independent model companies. He reiterated Z.AI's commitment to open-source models, framing it as part of the company's mission. He noted that even open-source models can engage in Other Agreements (OA) style collaborations.

Finally, the company highlighted a major achievement in domestic chip utilization. The GLM-5.3 Flash model, tested anonymously as “Ox-Alpha,” handled over 62 trillion tokens in just six days on OpenRouter, topping its charts. More importantly, it was the first model to serve ultra-large-scale real traffic entirely on a domestic chip cluster. Z.AI has achieved 100,000-card scale inference on domestic chips, reducing unit token costs by 80% since the beginning of the year. Chairman Liu Debin stated, "We are now beginning to control the cost curve ourselves." Looking ahead, Xiao Lei noted the next phase focuses not on whether domestic chips can run, but on validating their economic viability, anticipating a supply increase and improvement in heterogeneous computing environments in the next three to six months. The market reaction on the day of the earnings release was generally stable, with Z.AI's share price seeing a slight increase.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10