MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash
Fine, pp is still slower, but I'm slowly coming around to the idea of MTP finally being useful on Apple Silicon, and this is the first time I'm seeing a model outperform ds4 (and that's with IngeniousIdiocy's M3U tuning). MTP seems to have no advantage there as was always the case with llama.cpp, until now it seems. Qwen38FN will be the real test: vanilla ds4 currently spludging out 65 t/s (75 concurrently)...
评论
?
参与讨论