DI Series: Scaling GLM-5.1-FP8 to 64 MI300X GPUs

Serving a frontier Mixture-of-Experts (MoE) model well is a systems problem, and it gets harder the moment one node is not enough. GLM-5.1 is a good example: it is a large, sparse MoE that users want to run at long context, and it ships a new attention family that breaks assumptions older serving stacks quietly relied on. Fitting it on eight GPUs is only the start. The real question is how to keep it correct and fast as you spread it across several nodes.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论