Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?
Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling: AA says Qwen 3.8 27B is better by miles: artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b While Arena says Gemma 4 31B is almost 20 places ahead and completely trounces Qwen in many categories: arena.ai/leaderboard/text/overall The sentiment in this sub definitely seems in favour of Qwen, although not necessarily against Gemma which I think is still considered a good model. I recall poeple saying Qwen tends to be more tenacious and better at reasoning although at the cost of overthinking simple things. What is your explanation or experience with these models?