This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
They are high end really expensive Huawei ascend GPUs.
It is kinda bruteforcing the performance on a older semiconductor processing tech, so total production is pretty low.
By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
https://z.ai/blog/glm-5.3-flash