Does anyone actually respect benchmarks?
I get why they exist and in almost mostly any other hardware field we can see clearly the difference and what it respects throughout, but with ai, its so inconsistent and unpredictable, besides the very basic needle tests, which at this point what really fails it? I just dont get the hype around the benchmarks, ive been testing models that fit between 1-48gb vram for years now, everytime i go off a benchmark im usually disappointed, testing on my own workloads and env are the only sound testing i find shows anything actually useful for me I dont think anyone should worry about benchmarks so much when choosing a model, i know people consistently use the benchmarks to say z is better than y but you honestly need to test to see for yourself, unless you are talking a 9b model from 2 years ago vs a 27b released today, it might be hard to be certain what model is specifically best for yourself That being said, qwen has been the goat, and even after allmmy testing i seem to always stay/go back to their models, 35b + 3.8 27b right now are the best combo for speed/dense at my resources Curious if anyone else really feels this way or people actually respect these, useless benchmarks imho