Inside the ‘robot gyms’ training machines for the real world

Robotics labs are spending billions of dollars to generate the real-world data needed to train their machines, in a bid to overcome a deficit that is holding back “physical AI” start-ups’ development.

Unlike AI model makers, who scraped large parts of the web to develop their systems, a lack of training material has emerged as a serious bottleneck for robotics companies, which require a more diverse and complex dataset than text-based models.

Companies such as Generalist, Physical Intelligence, Neura and Skild are tapping external data providers and building out their own so-called robot gyms kitted out with bimanual manipulators, a set of robotic arms that workers operate and use to complete physical tasks that can later be emulated by robots.

Some start-ups are also shipping recording equipment to thousands of factories, homes, offices and holiday rentals all over the world to capture first-person video footage of people completing tasks, collecting millions of hours of this “egocentric” data.

“There is no internet for us to download from. There is no large corpus for model developers to work with,” said Joe Fox Jr, director of robotics operations at Scale AI, a data provider. “It is an arms race to get the best data.”

Robotics companies and physical AI labs are stepping up these costly projects after raising nearly $48bn in the year-to-date, according to PitchBook data, making it one of the hottest new corners of the AI boom for venture capital investors. A large portion of this figure will be spent on the equipment and tools used to generate data.

Data companies are feeding off physical AI fervour with start-up Mecka this week announcing it raised $60mn from investors, including Nvidia and Sequoia.

Ken Goldberg, a professor at UC Berkeley and co-founder of Ambi Robotics, said that while large language models had proven capable of achieving “remarkable things” using data scraped from the web, this was insufficient for robotics.

“It’s a completely new animal that needs to be trained from scratch with scarce data,” he said.

Physical Intelligence, a start-up that raised $600mn at a $5.6bn valuation last November, has tested its robotics in several Airbnbs across San Francisco in a bid to assess the capabilities of its models to handle different environments.

Figure, a humanoid robotics company, outlined in August a scheme to pay people to record themselves doing tasks in their home. It has roughly 44,000 active users recording themselves carrying out tasks in their home and workplace, who are paid an average of about 94 cents per task.

The Silicon Valley-based company plans to spend more than $1bn in 12 months on services and computing power to generate training data.

Encord, Scale AI and several other data companies that built a reputation for servicing AI labs such as OpenAI and Anthropic with annotated data to train large language models are opening up robot gyms or “data factories” around the world in the US, Mexico and parts of Asia, including China and Indonesia.

Robotics requires a more diverse set of data than the large language models that power Google’s Gemini, OpenAI’s ChatGPT and Anthropic’s Claude. Building a physical AI system requires video footage and hardware that can be used to generate data that can perceive depth and accurately calculate the amount of force required to carry out a particular task.

Robotics labs have also sought to use simulation, including so-called world models, to help bridge the gap in data.

A person operates two handheld controllers to guide robotic arms in folding a piece of clothing on a table.
Scale AI and several other data companies are opening up robot gyms or ‘data factories’ around the world © ScaleAI

Goldberg at Berkeley and his team are exploring the use of AI coding tools to help robots improve. Several companies also reported advances with “in-context learning” where a task is demonstrated by a person, which a robot then replicates.

Data labelling companies, which turned to cheaper labour markets to tag data for AI models, are continuing to do so in the robotics domain.

Scale has opened a data factory in Mexico City, while Xdof, which has raised $70mn from venture capital firms, including Andreessen Horowitz and Thrive Capital, has facilities in the Chinese city of Shenzhen and in Indonesia’s capital Jakarta.

Robotics operators are paid roughly $25 per hour in the US, nearly double what they are paid in Mexico, according to Glassdoor data and job listings for physical AI labs.

Philipp Wu, Xdof’s chief executive, said it was not economical to carry out work in the US given the high cost of labour and the volume of data required by physical AI labs.

Recommended

“The ideal situation is you maximise diversity . . . collecting [data] in all types of places,” Wu said. “So that when a robot is put in a new situation, it can generalise.”

Ulrik Stig Hansen, Encord’s co-founder and president, said that physical AI labs were using third parties to generate large volumes of data. But before contracting these data suppliers, robotics labs must first carry out experiments in-house to figure out whether a particular task or type of data would be useful if replicated multiple times.

Some operators are tapping factories in the global south in places such as Gurugram, India, where workers are being tasked with wearing recording equipment often without additional remuneration.

“We’re talking about filming in literal sweatshops,” said one executive at a data labelling company. “Even if the practice weren’t so gross, I question if the data itself is useful.”

Some robotics companies said they were circumspect about using third-party data. “The data that is most useful for us is the data we collect ourselves,” said one executive at a leading physical AI lab in San Francisco. “We actually do almost everything in-house because getting it right is really important.”

Companies hoping to replicate the “scaling laws” that have propelled advances in LLMs. That is, the more data the robotics models ingest, the better the output. Dyna, a robotics start-up, in a recent paper said that it had observed improved outputs in its model when training its models using up to 1mn hours of egocentric footage.

The sheer volume of data required means that spending by physical AI labs is expected to persist for some time.

“There’s no guarantee that if you generate a lot of data it will pay off,” Goldberg at Berkeley added. “Right now it’s a hunch and many people are willing to pay for it.”

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论