GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.

We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the optimal config is not the obvious one. Sharing the analysis because most of it applies wherever you rent or rack your B200s. The model GLM-5.2: ~750B total / ~40B active MoE (256 experts, top-8 routing, ~5.9% sparsity), DSA + MLA attention, 1M context, MIT license. Weights: ~744 GB in FP8, ~459 GB in NVFP4 (KV cache stays FP8). The hardware math 8x B200 SXM = 1,4

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论