In a quiet, nondescript warehouse in San Leandro, California, the future of artificial intelligence is not being written in lines of code, but in the deliberate, shaky removal of wooden blocks from a Jenga tower. Here, human "pilots"—specialized robotic trainers—meticulously perform physical tasks while tethered to an array of high-tech sensors. They are not playing games; they are generating the raw, high-fidelity data required to teach machines how to interact with the messy, unpredictable physical world.

This facility is the command center for Encord, a company that has pivoted from a mere data-management platform to a specialized factory for physical AI training data. As the robotics industry faces a critical bottleneck, Encord’s mission is clear: if the data needed to train a robot doesn’t exist, you have to manufacture it.

The Data Bottleneck: Why LLMs Don’t Translate to Hardware

For years, the AI narrative was dominated by Large Language Models (LLMs) that scraped the internet for text, images, and code. This "low-hanging fruit" allowed models like GPT-4 to achieve superhuman proficiency in digital tasks. However, the attempt to translate this success into physical robotics has hit a wall.

Vineeth Velmurugan, Encord’s head of robot learning and a veteran of OpenAI’s robotics lab and Berkshire Grey, notes that the "internet-scale" approach is insufficient for the physical world. "The data simply does not exist," he says. Unlike text, which can be scraped for pennies, physical data—video of a human hand navigating a server rack or stacking poker chips—is incredibly expensive and difficult to capture.

Velmurugan estimates that to reach a breakthrough in physical AI, the industry requires a dataset roughly five times the size of the entire YouTube video corpus. This massive scale requirement has transformed data collection from a secondary research problem into a primary business imperative.

Chronology of a Training Session: From Brainwaves to Robotic Limbs

The process at Encord’s San Leandro facility is a multi-modal ballet of human effort and sensor technology.

The Brain-Computer Interface (BCI) Trial

In a recent trial, Encord partnered with Zander Labs, a German neuroscience startup, to push the boundaries of data quality. A pilot, Andrew Ceja, wears a specialized headset designed to track brain activity. As Ceja dismantles the Jenga tower, the headset logs his "mental state"—detecting moments of surprise, frustration, or focused intent.

Lucas Gehrke, a neuroscientist at Zander Labs, suggests that this data is the "holy grail" for robot learning. By tagging training data with these cognitive markers, engineers can identify exactly when a model should deploy its most computationally expensive processes. If the human pilot is struggling or surprised, the robot learning model can prioritize that segment of data as a high-value teaching moment.

Egocentric Video and Muscle Sensing

Beyond neuro-sensors, Encord utilizes "egocentric" video—footage captured from the pilot’s point of view. To refine this, they are experimenting with electromyography (EMG) sensors strapped to the pilot’s forearms. These sensors detect electrical signals in human muscles, which Encord uses to map 3D hand positioning. Because cameras often struggle to capture the full dexterity of a hand during complex manipulations, these sensors provide a structural backbone that helps the robot understand depth and force in ways that 2D video cannot.

The Economics of "Manufactured" Data

The shift from managing data to manufacturing data has profound economic implications. For years, AI startups relied on crowdsourced, "junky" ego-data—randomized videos of human activity that are often poorly annotated.

Velmurugan argues that "dense annotation"—where every action, such as "right hand tightens bolt," is tagged with precise physical context—is roughly 100 times more valuable than unrefined data. While this gold-standard data costs roughly 20 times more to produce, the performance gains for a humanoid robot justify the expense.

However, this reality shatters the "LLM comparison." LLM makers built their empires on the back of free, publicly available data. Robotics companies, conversely, are entering a world of high-cost labor, bespoke hardware, and intense physical labor. This creates a "data moat" where only companies with the capital to fund these physical factories will likely emerge as winners in the humanoid race.

Industry Implications: The Vantage Point of the Middleman

Encord’s unique position as a service provider to various robotics firms gives them an unparalleled bird’s-eye view of the industry. Because they work with multiple undisclosed, high-profile robotics companies, they can observe, in real-time, which training techniques yield results and which are dead ends.

This vantage point is a significant strategic asset. By spotting trends—such as the industry-wide demand for precise manipulation of small objects like ethernet cables or kitchenware—Encord can optimize its facility to produce the most in-demand data sets before a single client even requests them.

The Human Element

The work is labor-intensive and requires a unique skill set. Pilots like Sofia Infante and Andrew Ceja, many of whom have backgrounds in previous AI annotation firms, must be part athlete, part technician. Their job is to bridge the gap between human intuition and machine limitation.

During a demonstration, Infante attempted to plug an ethernet cable into a server rack using robotic grippers. The struggle was immediate: human fingers have dozens of degrees of freedom, whereas robotic pincers are notoriously clumsy. This exercise highlights the fundamental challenge: even with perfect data, the hardware currently lags behind human dexterity. The pilots are effectively training the robots to overcome the mechanical limitations of the very bodies they inhabit.

Challenges and Future Outlook

Despite the progress, the "Physical AI" dream is still in its infancy. Several critical hurdles remain:

  1. Hardware Scalability: While data can be manufactured, the hardware (the robots themselves) remains expensive and fragile. Scaling this beyond a few San Leandro warehouses to a global level is a logistical nightmare.
  2. The "High-Effort" Problem: Even with brain-wave tagging, we do not yet know if training a robot on "human effort" levels will result in a model that is efficient in a production environment.
  3. Data Quality vs. Quantity: There is an ongoing debate within the field regarding whether we need more data or better data. While Velmurugan advocates for the "five times YouTube" scale, some researchers argue that high-quality, synthetic, or physics-simulated data could eventually replace the need for this expensive manual labor.

Conclusion: The New Industrial Revolution

Encord’s facility is a microcosm of a broader shift in the AI industry. We are moving away from the "infinite text" era into the "finite physical" era. The companies that win the next decade will not necessarily be those with the cleverest algorithms, but those with the most efficient "data manufacturing" pipelines.

As Ceja resets the Jenga tower for the tenth time, he remains unfazed by the repetition. For him, the constant flux of challenges is the job. For the robotics industry, however, that tower is a symbol of everything that remains to be solved. If they can successfully teach robots to handle the fragile, the sloshy, and the precise, they will have done more than just build a machine; they will have fundamentally changed the nature of physical labor in the modern world.

The race to build a functional, autonomous humanoid is no longer about who can hire the best software engineers. It is about who can build the best, most rigorous, and most insightful factories for human-to-robot knowledge transfer. In that race, Encord is, quite literally, holding the blocks.