Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?

Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling: AA says Qwen 3.8 27B is better by miles: artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b While Arena says Gemma 4 31B is almost 20 places ahead and completely trounces Qwen in many categories: arena.ai/leaderboard/text/overall The sentiment in this sub definitely seems in favour of Qwen, although not necessarily against Gemma which I think is still considered a good model. I recall poeple saying Qwen tends to be more tenacious and better at reasoning although at the cost of overthinking simple things. What is your explanation or experience with these models?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论