OTHER

Can Brainwaves Influence the Future of Physical AI?

Visualize a Jenga tower within a warehouse in San Leandro, California, representing the pinnacle of innovations in physical AI.

This location functions as the primary hub for Encord, a firm committed to creating data tools for training artificial intelligence models. Andrew Ceja, a pilot—Encord’s term for robotic trainers—carefully removes blocks from the delicate tower while donning a headset designed to capture his visual experiences. Although this approach is conventional, it integrates additional sensors to monitor his brain activity as he meticulously disassembles the tower.

Encord distinguishes itself from its emerging competitors by emphasizing that the primary obstacle for humanoid and warehouse robots is not the design of the models but the lack of authentic physical training data. Rather than merely providing existing datasets to robotics firms, Encord concentrates on generating the vital data that is currently missing.

The brainwave headset Ceja is using is created by Zander Labs, a neuroscience startup from Germany. They suggest that by analyzing brain activity—capturing cognitive elements such as error detection, intention, and surprise—they can produce more comprehensive datasets crucial for model training. Encord is working alongside Zander on a project to compile an initial dataset annotated with brainwave information, which will be tested against client robotic models to evaluate its effect on performance before broader application.

Lucas Gehrke, a neuroscientist at Zander spearheading this project, is hopeful that insights derived from recorded brain activity during different tasks can offer valuable guidance to model developers about when to implement their most advanced models.

Vineeth Velmurugan, Encord’s head of robotic learning, perceives this initiative as a groundbreaking solution to the data bottleneck in robotics. With experience from OpenAI’s robotics lab and Berkshire Grey, a company specializing in warehouse automation, Velmurugan joined Encord to build its internal data generation team.

Originally, Encord was established to assist organizations engaged in machine vision through data annotation and model assessment. However, they swiftly recognized a rising demand for their own training data as clients sought all-encompassing solutions for robotic manipulation tasks. “The data simply does not exist,” asserts Velmurugan.

The idea that generative AI could mirror its successes in chatbots for robotics faces a similar challenge. Large Language Models (LLMs) were trained on a vast range of online text, but acquiring comparable foundational data for training neural networks in physical manipulation is tough: companies working on self-driving automobiles typically gather their own data, but scaling this process can be tedious. While video data might serve as a potential training resource, it often lacks the precision found in real-world datasets. Velmurugan underscores that achieving groundbreaking results would necessitate a dataset approximately five times larger than the complete YouTube database, signifying that data generation poses a business challenge, not merely a research one.

Meeting Your Egocentric Data Needs

Currently, companies developing robotic intelligence predominantly rely on two main data sources: “Egocentric” video recordings captured by individuals with cameras, often mixed with various angles and metrics, and data gathered from remote-controlled robots. Encord employs both methods, sourcing egocentric data from factories globally while using its San Leandro facility to investigate novel techniques such as brainwave data and dataset generation for skill acquisition.

During a recent visit from TechCrunch, pilots utilized leader-follower configurations—paired robotic arms where one is controlled by a human and the other replicates its movements—to collect data for tasks like pouring coffee into mugs (a sophisticated task) and stacking poker chips. “Every humanoid robotics company is interested in these skills,” observed Velmurugan.

Storage areas were filled with boxes containing artificial flowers in vases, books, plastic vegetables, kitty litter trays, scoops, and bundles of wires—the essential tools required to train robots for household tasks.

At one station, pilot Sofia Infante adeptly maneuvers robotic arms to connect and disconnect ethernet cables from the back of a server—an operation that data center managers would be eager to automate if robots could achieve the necessary precision in cable management. After attempting to navigate the setup myself, I quickly realized why this task remains challenging: robotic claws lack the flexibility of human fingers and the range of motion we often take for granted.

Encord is also investigating a new data modality using sensors attached to the forearm to detect electrical signals in muscles. Since human hand movements are frequently not captured accurately on video, Velmurugan aims to construct a 3D model of hand positions at any given moment using arm sensors, which will contribute to a more complete dataset for model training.

Encord’s datasets are diligently annotated with detailed descriptions of actions depicted in each video—such as “right hand tightens bolt”—to assist LLM-based models in comprehending activities. Velmurugan estimates that this level of annotation detail is worth 100 times that of “low-quality ego data” for specific task training, while being only 20 times more costly to produce, making it initially appear advantageous.

However, being “20 times more” still signifies considerable costs, revealing the challenge: extracting text from the Internet—a method used by LLM developers who gather data from sources like Stack Overflow—incurs minimal expenses for large labs. In contrast, creating physical training data is not economically feasible, limiting the capacity of physical AI frameworks to closely replicate LLMs. This type of data must be generated rather than simply collected, thereby shifting the economic landscape of model development.

Velmurugan is optimistic that progress is being made—thanks to Encord’s insights into various industry initiatives; he observes that both startups and established laboratories are adopting strategies that enhance physical AI models. This unique perspective, connecting multiple robotics firms, strengthens Encord’s value proposition by proactively identifying emerging data methodologies ahead of individual clients.

This will keep Encord’s team of about a dozen pilots actively engaged. Both Infante and Ceja are part of a growing workforce dedicated to laying the foundation for neural networks; they transitioned from Scale, another AI data annotation company, to Encord.

Ceja previously worked in waste management, where his passion for technology allowed him to oversee a robotic waste sorting system. Now, while working with the Jenga tower, he conveys his excitement about tackling robot training challenges, stating, “It’s something new every day!”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.