Qwen3.6-27B - Effect of KV quantization on KLD - Q8, Q6, Q5 (bartowski)
Lower is better - Quantization increases from right to left I recently made a post here about how I squeezed more context into a Q8 model of bartowski's Qwen3.6-27B. My reasoning was that in my (anecdotal) experience, a Q8 has been performing a lot better than a Q6 or a Q5. There were a lot of comments about quantizing KV of a higher model and some folks suggested just going with a lower quant like Q6 but with full unquantized KV. So I just wanted to test that hypothesis with KLD. Base reference is Q8 with
评论
?
参与讨论