The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California, where Encord, a company that builds data tooling used to train AI models, is pushing the boundaries of robot learning.
Andrew Ceja, a pilot, carefully disassembles a wooden block tower while wearing a headset with a camera that tracks what he sees. This is fairly common for collecting robot training data, but the headset includes sensors that measure his brain waves as he works.
Measuring Brain Activity
The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup that’s betting measuring brain activity can create a more useful data set to train models. Encord’s work with Zander is currently a trial run, aiming to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up.
Lucas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models.
The Bleeding Edge of Robotics Data
This is the “bleeding edge” of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord’s head of robot learning. A veteran of OpenAI’s robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company’s internal data-creation team.
Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. “The data simply does not exist,” Velmurugan said.
The bet that generative AI can do for robots what it’s done for chatbots keeps running into this same wall. LLMs were built on the text of the entire internet, and more. Finding the same raw materials to teach neural networks about physical manipulation is challenging: self-driving car companies collect it themselves, but that’s hard to scale. Training from video can work, but it lacks the fidelity of real-world data. Velmurugan says it will take a data set something like five times the size of YouTube’s video corpus to break through—a scale that helps explain why data-generation itself has become a business and not just a research problem.
Generating Physical Training Data
Companies building robot brains are now turning to two main sources: “Egocentric” video collected by workers wearing cameras, often augmented with additional camera angles and other metrics, and collecting data from robots operated remotely. Encord does both, drawing egocentric data from several factories around the globe, and using its San Leandro facility to experiment with new modalities, like brain waves, or collect data sets around specific skills for fine-tuning.
When TechCrunch visited, pilots were using leader-follower rigs — paired robotic arms, one controlled directly by a human operator and one that mimics its movements —to create data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan says.
Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires, the stock in trade for training manipulators for household tasks.
At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to plug and unplug ethernet cables from the back of a server—the kind of work data center operators would love to be automated, if only robots could manipulate them with the required precision. Taking a spin behind the controls, I was able to see why that’s still out of reach: Pincers are far less dextrous than human fingers and lack the degrees of freedom we take for granted in our arms.
Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically doesn’t capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models.
Encord’s data sets are annotated with physical descriptions of what each video contains—”right hand tightens bolt”—to aid LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as “junky ego data” for training specific tasks, and it only costs 20 times more to produce, which is a good trade, on paper.
But “20 times more” is still real money, and that’s the catch: scraping text off the internet, the way LLM makers built their models by pulling from Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not, and that’s the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models.
Breaking Through the Data Bottleneck
Velmurugan says that progress is being made—with Encord’s visibility into programs across the industry, he’s able to see start-ups and frontier labs alike figure out what works and what doesn’t to improve physical AI models. That vantage point—sitting between many robotics companies at once—is also part of Encord’s pitch. It can spot which data techniques are gaining traction industry-wide before any single customer can.