Qwen3.8-27B thinking xhigh Vs. thinking off - Apple M5 Max

I ran a few benchmarks on my MacBook Pro M5 Max with oMLX: Running Qwen3.8-27B with thinking off has a big effect on output quality. Running it with thinking on xhigh burns 5.5x the tokens and runs 6x longer. The amount of thinking that Qwen3.8-27B puts into its work is really enormous - this has already been discussed a lot. The big impact of turning thinking completely off is huge - quality wise it basically puts the model behind Qwen3.6-35B-A3B and other MoE models that provide way more Tok/s. Details here: Compare Benchmark Runs | llm-bench.io

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论