Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params

Weights, all 4 sizes, Apache-2.0 (ViT-S / ViT-B / ViT-L / ViT-g): huggingface.co/collections/robbyant/lingbot-vision Code: github.com/robbyant/lingbot-vision Project page: technology.robbyant.com/lingbot-vision Self-supervised DINO-family backbone, but the masking is boundary-driven: the teacher predicts where object boundaries are and those tokens get forced into the student's mask, so it can't solve reconstruction by copying flat context. No labels, no text supervision, no external

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论