I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)

I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results are interesting. Of course, the smaller file gives worse results. However they are not that far off. Unfortunately, this comes at the expense of even more tokens beeing used by the Bonsai model and thus much longer generation times. Visually I prefer the IQ3_XXS results, but see for yourself. The test is by no means scientific - just few UI generation tasks for direct comparison on the same hardware. Also, I ran llama.cpp with MTP while the Bonsai model doesn't seem to have MTP which makes it even slower. Metric Qwen IQ3 Bonsai PQ2 Tasks completed 4/4 4/4 Fixed assertions 20/20 20/20 Agent wall time 8:00 24:09 Output tokens 27,197 84,176 Weighted decode 83.59 tok/s 64.91 tok/s Speculative acceptance 65.22% MTP 39.76% modified N-gram Compactions 0 0 Length stops 0 1 Across the complete suite, Qwen finished 3.02× faster and used 3.10× fewer output tokens . Results: danmoreng.github.io/qwen3-8-27b-iq3-xxs-vs-bonsai Repo: github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论