mistral.rs v0.9.0: up to 1.8x faster CPU decode than llama.cpp on x86 and ARM!

preview.redd.it/nuk5rxceptbh1.png On Qwen3 4B Q4_K, mistral.rs decodes faster than llama.cpp at every context depth we measured, on x86 (Sapphire Rapids) and ARM (GB10). We optimized mistral.rs at granular levels to achieve general speedups for all models . Additionally, our optimizations apply to CPUs of all calibers : from x86 with AVX2 or AVX512, to ARM processors with NEON. We wanted to make sure this was a comparison in

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论