Why are MoE models so belittled?

E.g "Qwen 3.5 122B is just 10B active, so it's no where close to the dense 27B model" That is the main sentiment around here and it puzzles me. If a 122B is just worth 10B, then why does model providers bother creating an MoE model when they could've just released a dense 10B model? Heck the 10B dense would run faster than the 122B MoE (no routing overhead), which negates the supposed ( only advantage of MoE is speed ) argument. It sure is not that simple. I mean yes it's only 10B active at a time, but it c

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论