pareto frontiers
Qwen models are both pareto frontiers in total size AND active parameters size of all open weights models so far. If this trend continues, we might see sparser and more capable models really soon, given that this is a preview of Qwen4 and is probably undertrained. What do you people think? Will the trend continue? Will Qwen stay in the lead? And more importantly, does it scale up? (e.g. would Qwen4 architecture at larger scales be even better? or diminishing returns?) I personally like the path towards more sparse, fast and capable models. n-grams really are a step change for local AI. And hopefully, prices of hardware will go down or be affordable enough to run such models, alongside software improvements to run the best possible intelligence on existing hardware. I'd love to hear everyone's thoughts and/or differing opinions and arguments, constructively. Source: AA