Can Brainwaves Influence the Evolution of Physical AI?
Visualize a Jenga tower located in a warehouse in San Leandro, California, representing the cutting edge of physical AI technology.
This facility is the headquarters of Encord, a company focused on creating data tools for training AI models. Andrew Ceja, referred to as a pilot—Encord’s designation for robotic trainers—meticulously removes blocks from the delicate tower while wearing a headset that captures his visual experience. While conventional, this data-gathering method includes additional sensors to monitor his brain activity as he carefully takes apart the tower.
Encord distinguishes itself from new competitors by asserting that the most significant issue for humanoid and warehouse robots lies not in the design of the models, but in the lack of real-world physical training data. Rather than just providing existing datasets to robotics firms, Encord is committed to producing the essential data that is currently missing.
The brainwave headset Ceja uses is developed by Zander Labs, a neuroscience startup from Germany. They argue that analyzing brain activity—capturing mental states such as error detection, intent, and surprise—can yield richer datasets essential for model training. Encord and Zander are currently collaborating on a project to create an initial dataset annotated with brainwave data, which will be tested with client robotic models to evaluate its impact on performance prior to fully adopting this method.
Lucas Gehrke, a neuroscientist from Zander overseeing this initiative, is hopeful that insights gained from recorded brain activity during various tasks can provide valuable direction to model developers on the optimal timing for integrating their most advanced models.
Vineeth Velmurugan, Encord’s head of robotic learning, views this effort as a groundbreaking attempt to tackle the data bottleneck in robotics. After his experience at OpenAI’s robotics lab and Berkshire Grey, a company specializing in warehouse automation, Velmurugan joined Encord to establish its internal data generation team.
Originally, Encord was founded to assist organizations involved in machine vision by offering data annotation and model evaluation. However, they quickly recognized the urgent need to produce their own training data as clients demanded comprehensive solutions for robotic manipulation tasks. “The data simply does not exist,” Velmurugan emphasizes.
The idea that generative AI might replicate its chatbot successes in robotics faces a similar obstacle. Large Language Models (LLMs) were trained on a diverse range of texts from the Internet and beyond. However, discovering comparable foundational resources for training neural networks focused on physical manipulation poses a challenge: companies making self-driving cars typically gather their own data, but scaling this process can be exhausting. Although video data offers a potential source for training, it often lacks the accuracy of real-world datasets. Velmurugan argues that achieving transformative results would require a dataset approximately five times larger than the entire YouTube library, underscoring that data generation is primarily a business challenge, not merely a research issue.
Meeting Your Egocentric Data Needs
Currently, organizations developing robotic intelligences mainly focus on two primary data sources: “Egocentric” video footage recorded by individuals using cameras, often complemented with various angles and metrics, in addition to data gathered from remote-controlled robots. Encord employs both strategies, amassing egocentric data from numerous factories globally while utilizing its San Leandro facility to explore innovative methods, such as brainwave data and dataset generation for skill acquisition.
During a recent visit from TechCrunch, pilots were implementing leader-follower configurations—paired robotic arms, where one is operated by a human and the other mirrors its actions—to gather data for tasks like pouring coffee into mugs (a challenging endeavor) and stacking poker chips. “Every humanoid robotics company has shown interest in these skills,” Velmurugan noted.
Storage areas were stocked with boxes of artificial flowers in vases, books, plastic vegetables, kitty litter trays, scoops, and bundles of wires—the crucial tools for training robots to complete household tasks.
At one station, another pilot, Sofia Infante, skillfully manipulates robotic arms to connect and disconnect ethernet cables from a server’s rear—an operation that data center operators would eagerly automate if robots could manage cables with the precision required. After attempting to navigate the setup myself, I quickly understood why this task remains challenging: robotic claws are far less flexible than human fingers and lack the range of motion we often take for granted.
Encord is also investigating a new data modality using sensors attached to the forearm to detect electrical signals in muscles. Human hand movements are often inadequately captured on video, but Velmurugan aims to construct a 3D model of hand positions at any moment utilizing arm sensors, resulting in a more comprehensive dataset for model training.
Encord’s datasets are carefully annotated with detailed descriptions of actions depicted in each video—such as “right hand tightens bolt”—to assist LLM-based models in understanding activities. Velmurugan estimates that this detailed level of annotation is valued at 100 times that of “low-quality ego data” for specific task training, while costing only 20 times more to produce, making it seemingly advantageous at first glance.
However, “20 times more” still signifies substantial costs, highlighting the challenge: extracting text from the Internet—a method employed by LLM developers who gather data from platforms like Stack Overflow—incurs minimal expenses for large laboratories. Conversely, generating physical training data is not economically feasible, constraining the ability of physical AI frameworks to closely mirror LLMs. This type of data must be created rather than simply collected, altering the economic landscape of model development.
Velmurugan asserts that progress is being made—thanks to Encord’s insights into various industry initiatives, he observes that both startups and established labs are implementing strategies that enhance physical AI models. This unique perspective, bridging multiple robotics firms, amplifies Encord’s value proposition by proactively identifying emerging data methodologies before any individual client.
This will keep Encord’s team of approximately a dozen pilots fully engaged. Both Infante and Ceja are part of a growing workforce establishing the foundation for neural networks; they transitioned from Scale, another AI data annotation company, to Encord.
Ceja previously worked in waste management, where his passion for technology led him to oversee a robotic waste sorting system. Now, while interacting with the Jenga tower, he expresses enthusiasm for tackling robot training challenges, stating, “It’s something new every day!”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


