Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp

TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to ~400. Dropping DCP raises prefil to ~750. I'm running 4 drafted tokens vs Z.ai's rec of 5. Decode is heavily dependent on prose. Thinking gets ~20 tok/s. Code gets 25-35. Typical turns in Pi get me ~24 tok/s. Pruning the model by 5-10% will probably get you to 1M ctx or more concurrency if you need that. In my daily use, a 10% data-free prune seems to preserve the m

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论