DVD-JEPA: an open-source, fully-reproducible JEPA world model [P]

A paper currently trending on paperswithcode.co in the "Anomaly Detection" category is DVD-JEPA . i.redd.it/r6fd8n3d4f8h1.gif Here is the short summary: Most attempts to learn a world model from video try to predict the next frame pixel-by-pixel, and drown in detail that is fundamentally unpredictable. JEPA (Joint-Embedding Predictive Architecture, LeCun 2022 ) makes a different bet: predict the representation of the future, not the pixels, and let the encoder discard whatever it cannot predict. DVD

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论