I got Nemotron Puzzle 75B running smoothly on a 64GB M2 Max
TL;DR: Added native nemotron_h_puzzle support to mlx-lm ( PR #1535 ), then compared 4-bit vs 5-bit expert quantization (both with 6-bit dense layers, BF16 output head, group size 64) on a 64GB M2 Max. Results (same prompts, 5 seeds per task, temp 1.0 / top_p 0.95): 4-bit experts 5-bit experts Dense paths 6-bit 6-bit Output head BF16 BF16 Checkpoint 42.03 GiB 49.88 GiB Peak MLX memory 49.68 GB 58.12 GB Average generation 14.27 tok/s 10.53 tok/s Local task checks 24/30 21/30 Long-context retrieval 4/5 0/5 i b
评论
?
参与讨论