LocateAnything Explained: Parallel Box Decoding and the Next Generation of Vision-Language Grounding

🫧 Open on Bubbles A review of LocateAnything, an NVIDIA vision-language model that treats each bounding box as one atomic unit and decodes it in a single parallel step instead of a sequence of coordinate tokens. Its Parallel Box Decoding reaches roughly 2.5x the throughput of the nearest grounding VLM while improving high-IoU localization.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论