LLM Quantization Part 3: Honey, I Shrunk the Numbers!
By now we know what a model is made of, why it eats so much VRAM, and what those weights actually are: billions of numbers, most of them sitting in a tiny band near zero, stored in a format built to reach 3.4 × 10³⁸. Let us look at how shrinking the numbers actually works!
评论
?
参与讨论