💻 Z.ai built production inference for GLM-5.3-Flash on 100,000+ Chinese accelerators

The company published an article, “Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure”: a significant part of the engineering was done by an Infra Agent based on GLM-5.3, and it took less than two weeks from the first adaptation to production readiness.

🌍 According to Z.ai, end-to-end throughput increased 3-fold, and token cost and hardware utilization reached the level of mainstream NVIDIA cards. These are vendor figures without independent verification, but it is a precedent: a flagship model in production on 100k+ non-NVIDIA chips.

👤 Before its public launch, the model worked anonymously under the name Ox-Alpha on OpenCode and OpenRouter and became the most used on both platforms within a week, processing more than 62 trillion tokens in 6 days. Non-NVIDIA infrastructure is already handling real production loads.

Source 1: https://z.ai/blog/glm-built-its-inference-infrastructure