Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params
Weights, all 4 sizes, Apache-2.0 (ViT-S / ViT-B / ViT-L / ViT-g): huggingface.co/collections/robbyant/lingbot-vision Code: github.com/robbyant/lingbot-vision Project page: technology.robbyant.com/lingbot-vision Self-supervised DINO-family backbone, but the masking is boundary-driven: the teacher predicts where object boundaries are and those tokens get forced into the student's mask, so it can't solve reconstruction by copying flat context. No labels, no text supervision, no external
评论
?
参与讨论