In current ML systems, where is the main bottleneck: dataset quality or model architecture improvements? [D]

A lot of recent progress in ML appears to come from scaling existing architectures rather than introducing fundamentally new ones. At the same time, there’s increasing emphasis on dataset quality, curation, and synthetic data pipelines. In practice, I’m trying to understand how this tradeoff looks in real systems: How much effort is typically spent on data cleaning and filtering vs model design?? Whether dataset quality improvements still yield larger gains compared to architectural changes?? How synthetic

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论