4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
TL;DR: Full GLM-5.2 (753B MoE) quantized to Int4-Int8Mix + NVFP4 4-bit KV cache, TP=4 across 4× DGX Spark (GB10) at 100K context , run on Terminal-Bench 2.1 with the same agent scaffold (Terminus-2) as the official numbers. Result: 63/89 = 70.8% vs the official full-precision 81.0% . Caveat up front: I never ran the full model through my pipeline — the ~10-pt gap bundles quantization plus my 100K-vs-256K context cap, a smaller token budget, and unmatched sampling. So read it as: my whole 4-bit/100K desktop
评论
?
参与讨论