Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure
kept by eddie
GLM developed a production inference service on over 100,000 domestically made AI accelerators to run its GLM-5.3‑Flash model and achieved significant performance gains through aggressive memory optimizations.