Qwen3.8-27B MTP quants on Apple M5 Max — which one is actually worth it?
Ran a full oQ2e → oQ8e sweep yesterday. Here's what the numbers say: | Quant | Speed | Quality Score | |-----------|-------------|-------------------------| | oQ2e-mtp | 44.8 tok/s | 16.9 ❌ (unusable) | | oQ3e-mtp | 38.7 tok/s | 85.2 ✅ | | oQ4e-mtp | 36.6 tok/s | 86.8 ✅ best balance | | oQ6e-mtp | 29.0 tok/s | 85.7 ✅ | | oQ8e-mtp | 27.2 tok/s | 87.0 🏆 highest quality | ==> oQ2e is not usabel. ==> oQ3e already delivers good quality. ==> oQ4e is the sweet spot in terms of speed and quality. ==> oQ6e not a real gain in terms of quality over oQ4e but at 20% lower tok/s output. ==> oQ8e if you want the full quality range and can live with 25% lower tok/s than oQ4e and 2x the memory footprint Full benchmark results (all hardware, all quants): llm-bench.io Qwen3.8-27B MTP quant comparison