Gemini Robotics 2 brings us one step closer to physical AGI

Gemini Robotics 2 brings us one step closer to physical AGI 图片 1
Gemini Robotics 2 brings us one step closer to physical AGI 图片 2
Gemini Robotics 2 brings us one step closer to physical AGI 图片 3

This week, Google DeepMind revealed Gemini Robotics 2, an intelligence layer comprising three new models to power more adaptable physical AI. Together, Google says these models will give robots more dexterous, full-body control to work together and complete a wide range of multi-step tasks.

Google announced the vision-language-action (VLA) model, Gemini Robotics 2, on Thursday, and converts vision and language inputs into motor control so robots can flex from feet to fingertips with enough dexterity in hands and grippers to complete delicate tasks, like closing a Ziploc bag. The lightweight version, Gemini Robotics On-Device 2, runs locally so robotic applications can keep running even without internet connectivity.

Meanwhile, the embodied reasoning (ER) model, Gemini Robotics ER 2, lets robots understand their surroundings and communicate with humans so they can devise plans to carry out multi-step tasks — think emptying a dishwasher and putting items away.

Chris Matthieu, VP of the developer ecosystem at RealSense, tells The New Stack this is what it’ll take for robots to graduate from simple, isolated actions to real-world physical assistance:

“The hard part isn’t making the first decision — it’s recovering from the hundredth when the world has changed. Doors are closed, objects get moved, people walk into the scene, batteries drain, and sensors become partially occluded.”

Full-body control and greater dexterity

According to Google, combining VLA and ER models means humanoids can go further to complete a range of tasks, literally.

Where its previous Gemini Robotics 1.5 model could only control a robot’s upper body to perform tabletop tasks, Gemini Robotics 2 brings intelligence to the entire humanoid body, enabling it to walk, crouch, stretch, and handle various objects. That means developers could create systems where robots are capable of executing commands like, “fetch the book from the top shelf.”

“Humans make this look effortless because our visual…
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论