3D-Object Perception Transformer — CVPR 2026 Highlight — OpenCV Live! 222
The 3D-Object Perception Transformer (3PT) unifies detection, segmentation, and 6DoF pose estimation into two multi-view, RGB-only transformers. Demonstrating exceptional accuracy and cross-domain robustness, it placed first by significant margins in both the Industrial Robotics and AR/VR tracks of the BOP 2025 challenge at ICCV. Today, the 3PT architecture is actively deployed in real-world industrial robotic workcells across the world as the Intrinsic Vision Model (IVM).
Our guest this week, Agastya Kalra — (Intrinsic (Google), University of Hawaii) was an author of the paper and will walk through the architecture, the BOP 2025 results, and what it took to move the model from benchmark to production robotics.
Read the paper: https://www.intrinsic.ai/publications/3pt-cvpr2026
Join our live stream and stick around for our giveaway of a free OpenCV University course to one lucky viewer.
- Watch on Patreon (DRM-Free, Ad-Free!)
- Watch on YouTube
- Watch on Apple Podcasts
- Watch on LinkedIn
- Watch on Zoom
- Watch on Twitch
Got a cool project of your own? Send it to us and you may be featured on a future episode.
The post appeared first on OpenCV.