Skip to content
SemiAnalysis@SemiAnalysis_ · Aug 26, 2026

Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese…

3 tweets1 min read♥ 2.7Koriginal

Summary

Ox Alpha, revealed to be GLM-5.3-Flash, is serving 100 trillion tokens per day on Chinese chips rather than Nvidia GPUs, achieving comparable efficiency and cost-effectiveness. This development challenges Nvidia's CUDA dominance and shows that frontier-level AI compute can be delivered efficiently on alternative hardware.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese chip. (1/3)🧵

  2. #2

    100T tokens per day free tokens and people were saying only frontier labs has this amount of compute. (2/3)

    Theo - t3.gg@theo · Aug 21, 2026

    “We have capacity for 100T tokens per day” Okay who the fuck made this model and where did they get this much compute?

  3. #3

    But ALL traffic was served on Chinese chips, attaining hardware efficiency and per-token cost comparable to Nvidia GPUs. The cuda moat is being tested once again after Jalapeño's announcement yesterday. (3/3)