India's Motion Farms: Human Labor Is Still the Cheapest Way to Train Humanoid Robots
India’s factory workers are wearing head-mounted cameras so that robots can learn to use their hands.
The arrangement is practical, not symbolic. Humanoid robots — the kind being built by Figure, 1X Technologies, Physical Intelligence, and a dozen others racing toward commercial deployment — do not learn from the internet the way large language models did. Text at scale gave LLMs syntax, reasoning, and factual recall. It gives robots nothing useful about grip angle, slip correction, or how to fold a garment without tearing it.
That gap is the physical intelligence bottleneck. And for now, the cheapest way to fill it is to put cameras on workers and record what they already do.
What Gets Captured
The data isn’t just video. First-person recordings from camera-equipped workers capture the micro-structure of physical work: where a hand starts relative to an object, how force changes as fingers contact surface irregularities, when wrists rotate, how grip adjusts after a near-slip. That sequencing — posture, bimanual coordination, the micro-recovery after contact goes wrong — is the exact layer that’s hard to synthesise and expensive to collect from robot hardware.
Teleoperated robots require equipment, skilled operators, continuous calibration, and constant fail recovery. A factory worker in a packaging or sorting facility costs a fraction as much per hour of usable data and works in the exact high-contact environments — folding, grasping, tool use, object sorting — where first-generation humanoid deployment is targeted.
Why Text Doesn’t Scale to Physics
LLMs got lucky: their training domain (text) was already abundant, cheap, and machine-readable. Physical AI has no equivalent. The internet contains no ground-truth data about how a wrist should orient to lift a lid without spilling, or how fingers should sequence on a power drill to avoid cam-out.
Simulation helps at the margins — synthetic environments can generate action data faster than reality — but sim-to-real transfer degrades badly for contact-rich tasks involving deformable objects, liquids, and surface variation. Real-world motion data, even imperfect video, remains the reference distribution that simulated data has to approximate.
The Data Moat
For humanoid labs, proprietary motion datasets are becoming a competitive moat with the same structure as proprietary training data had for LLMs in 2022–23. Every new manipulation category — liquid containers, deformable materials, precision assembly, bi-manual coordination — requires a separate collection campaign.
Labs investing in structured physical data collection now are building an asset that compounds. The question is not which lab has the best model architecture. It is which lab has the highest-quality, widest-distribution embodied training data — and whether they got it before commodity collection pipelines (cheaper robot teleoperation, better sim transfer, synthetic generation) close the gap.
The Uncomfortable Math
The labor force most directly in the path of humanoid automation is also, at this moment, the cheapest source of data for training it. Until embodied data infrastructure scales — robot fleets large enough to self-generate experience, simulation pipelines robust enough to transfer, synthetic generation reliable enough for contact tasks — humanoid labs will keep extracting physical intelligence from the workers the robots are ultimately meant to replace.
That feedback loop is not unique to AI. But the speed of the transition, and the explicitness of the extraction — workers wearing sensors so their motions can be packaged as supervised learning targets — makes the dependency unusually visible.