Can Brainwaves Influence the Evolution of Physical AI?
Envision a Jenga tower standing in a San Leandro, California warehouse, representing the pinnacle of physical AI advancements.
This facility acts as the headquarters for Encord, a firm committed to creating data tools for training AI models. Andrew Ceja, known as a pilot—Encord’s term for robotic trainers—carefully removes blocks from the delicate tower while wearing a headset that records his visual experiences. Although this approach is conventional, it integrates extra sensors to monitor his brain activity as he meticulously disassembles the tower.
Encord differentiates itself from its emerging competitors by emphasizing that the key challenge for humanoid and warehouse robots isn’t their model design, but the lack of real-world physical training data. Rather than just providing existing datasets to robotics firms, Encord prioritizes the creation of essential data that is currently missing.
The brainwave headset worn by Ceja is developed by Zander Labs, a Germany-based neuroscience startup. They suggest that by analyzing brain activity—capturing cognitive facets like error detection, intention, and surprise—they can generate more comprehensive datasets vital for model training. Encord and Zander are collaborating on a project to produce an initial dataset annotated with brainwave information, which will be evaluated against client robotic models to measure its influence on performance before widespread use.
Lucas Gehrke, a Zander neuroscientist leading this project, is hopeful that insights from recorded brain activity during different tasks can offer valuable guidance to model developers regarding when to utilize their most sophisticated models.
Vineeth Velmurugan, Encord’s head of robotic learning, perceives this initiative as a groundbreaking move to tackle the data bottleneck in robotics. After working at OpenAI’s robotics lab and Berkshire Grey, a company specializing in warehouse automation, Velmurugan joined Encord to create its internal data generation team.
Initially, Encord was established to assist organizations engaged in machine vision through data annotation and model evaluation. However, they soon recognized a rising demand for their own training data as clients sought all-encompassing solutions for robotic manipulation tasks. “The data simply does not exist,” Velmurugan states.
The idea that generative AI might replicate its chatbot triumphs in robotics faces a similar challenge. Large Language Models (LLMs) were trained on an extensive collection of online text. However, acquiring comparable foundational resources for training neural networks focused on physical manipulation is complicated: companies developing self-driving cars typically assemble their data, but scaling this process can be labor-intensive. While video data could potentially serve as a training source, it often lacks the accuracy of real-world datasets. Velmurugan insists that achieving transformative results would require a dataset approximately five times larger than the entire YouTube database, emphasizing that data generation is a business challenge, not solely a research one.
Meeting Your Egocentric Data Needs
Currently, organizations creating robotic intelligences mainly depend on two primary data sources: “Egocentric” video footage recorded by individuals with cameras, often blended with various perspectives and metrics, and data sourced from remote-controlled robots. Encord employs both methods, gathering egocentric data from factories globally while using its San Leandro facility to explore innovative strategies such as brainwave data and dataset generation for skill acquisition.
During a recent visit from TechCrunch, pilots utilized leader-follower setups—paired robotic arms, where one is operated by a human and the other mirrors its movements—to collect data for tasks like pouring coffee into mugs (a challenging endeavor) and stacking poker chips. “Every humanoid robotics company has shown interest in these skills,” noted Velmurugan.
Storage areas were packed with boxes filled with artificial flowers in vases, books, plastic vegetables, kitty litter trays, scoops, and bundles of wires—the essential instruments required to train robots for domestic tasks.
At one station, another pilot, Sofia Infante, skillfully maneuvers robotic arms to connect and disconnect ethernet cables from the back of a server—an activity that data center managers would eagerly automate if robots could achieve the required precision in cable management. After attempting to navigate the setup myself, I quickly grasped why this task remains challenging: robotic claws are far less adaptable than human fingers and lack the range of motion we often take for granted.
Encord is also investigating a new data modality using sensors attached to the forearm to detect electrical signals in muscles. Human hand movements are often inadequately captured on video, but Velmurugan aims to construct a 3D model of hand positions at any given moment using arm sensors, leading to a more thorough dataset for model training.
Encord’s datasets are meticulously annotated with detailed descriptions of actions shown in each video—such as “right hand tightens bolt”—to aid LLM-based models in comprehending activities. Velmurugan estimates that this level of annotation detail is valued at 100 times that of “low-quality ego data” for specific task training, while costing only 20 times more to produce, making it seem advantageous initially.
However, “20 times more” still represents significant costs, highlighting the challenge: extracting text from the Internet—a strategy used by LLM developers who gather data from platforms like Stack Overflow—incurs minimal expenses for large labs. In contrast, generating physical training data is not economically feasible, limiting the capacity of physical AI frameworks to closely mimic LLMs. This form of data must be created, not simply collected, thereby reshaping the economic framework of model development.
Velmurugan contends that progress is being made—thanks to Encord’s insights into various industry initiatives, he observes that both startups and established labs are implementing strategies that enhance physical AI models. This distinct perspective, connecting multiple robotics firms, strengthens Encord’s value proposition by proactively identifying emerging data methodologies ahead of individual clients.
This will keep Encord’s team of about a dozen pilots fully engaged. Both Infante and Ceja are part of an expanding workforce laying the groundwork for neural networks; they transitioned from Scale, another AI data annotation company, to Encord.
Ceja previously worked in waste management, where his passion for technology enabled him to oversee a robotic waste sorting system. Now, while interacting with the Jenga tower, he conveys excitement about tackling robot training challenges, stating, “It’s something new every day!”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


