Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese…
Summary
Ox Alpha, revealed to be GLM-5.3-Flash, is serving 100 trillion tokens per day on Chinese chips rather than Nvidia GPUs, achieving comparable efficiency and cost-effectiveness. This development challenges Nvidia's CUDA dominance and shows that frontier-level AI compute can be delivered efficiently on alternative hardware.
Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.
- #1
- #2
- #3