Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to ~400. Dropping DCP raises prefil to ~750. I'm running 4 drafted tokens vs Z.ai's rec of 5. Decode is heavily dependent on prose. Thinking gets ~20 tok/s. Code gets 25-35. Typical turns in Pi get me ~24 tok/s. Pruning the model by 5-10% will probably get you to 1M ctx or more concurrency if you need that. In my daily use, a 10% data-free prune seems to preserve the m
评论
?
参与讨论