LocateAnything Explained: Parallel Box Decoding and the Next Generation of Vision-Language Grounding
🫧 Open on Bubbles A review of LocateAnything, an NVIDIA vision-language model that treats each bounding box as one atomic unit and decodes it in a single parallel step instead of a sequence of coordinate tokens. Its Parallel Box Decoding reaches roughly 2.5x the throughput of the nearest grounding VLM while improving high-IoU localization.
评论
?
参与讨论