LLM Quantization Part 3: Honey, I Shrunk the Numbers!

By now we know what a model is made of, why it eats so much VRAM, and what those weights actually are: billions of numbers, most of them sitting in a tiny band near zero, stored in a format built to reach 3.4 × 10³⁸. Let us look at how shrinking the numbers actually works!

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论