XDOF Raises $70M From a16z and Thrive Capital to Build the Data Infrastructure Layer for Physical AI
Every major humanoid robotics company — 1X, Figure, Agility, Boston Dynamics — shares the same bottleneck: the robots need physical interaction data at scale, and that data does not exist in the volumes that LLM training did. The internet did not produce a trillion robot manipulation sequences. Someone has to collect them.
XDOF is building the infrastructure to do that. The company emerged from stealth on June 17 with $70 million in backing from Thrive Capital, a16z, Spark Capital, Lux, and WndrCo. Its pitch is straightforward: it is the data layer for physical AI.
What XDOF Actually Builds
XDOF’s product stack has three components:
- Data pipelines: Ingestion and processing infrastructure for physical interaction recordings at production scale
- Collection tools: Hardware and software systems for capturing robot manipulation data, including egocentric video and teleoperation sequences
- Annotation systems: The labeling and cleaning layer that converts raw recordings into training-grade datasets
The company employs teleoperators — humans who control robots remotely to generate ground-truth motion data — and egocentric data operators who wear sensor rigs to capture first-person physical interaction from a human perspective. Both generate data types that robot foundation models currently lack at scale.
XDOF has also partnered with UC Berkeley to release a large collection of robot manipulation data to the research community. Berkeley’s involvement gives the dataset institutional credibility and provides a public-domain training corpus that smaller robotics labs can access without XDOF’s commercial tier.
The Scale AI Comparison
In the LLM era, the equivalent bottleneck was text and image labeling. Scale AI solved it by building industrial annotation infrastructure and became a billion-dollar company before most frontier models were public. XDOF is making the same bet one layer earlier in the stack: the physical world requires its own data collection infrastructure, and whoever builds the best pipelines earliest will be embedded in every major robotics company’s training workflow.
The parallel has limits. Text labeling is location-independent. Robot training data requires physical facilities, hardware, and operators in specific environments. XDOF’s cost structure is materially heavier than a software annotation platform.
The $70 million raise indicates its backers have done the math and concluded the market is large enough to absorb it. Global humanoid shipments rose 800% in 2025 according to recent industry data, with China fielding 140 manufacturers. Demand for training data is not ahead of demand for robots — it is concurrent with it.
Why This Round Matters
The investor list is notable. Thrive Capital and a16z both have direct exposure to physical AI through other portfolio companies — Thrive through robotics and XDOF gives them a position in the data layer that feeds the entire sector. Lux Capital has backed physical science bets consistently. Spark Capital’s presence rounds out a coalition that is not making a speculative bet on humanoids eventually arriving; it is making a supply chain bet on the infrastructure that robots already need right now.
WndrCo, the Jeffrey Katzenberg vehicle, adds an unusual entry: media and entertainment have significant interest in robot performance and controlled motion, suggesting XDOF’s teleoperation data may have value outside pure manufacturing robotics.
Key Numbers
- Raise: $70M from stealth
- Investors: Thrive Capital, a16z, Spark Capital, Lux, WndrCo
- UC Berkeley partnership: large-scale robot manipulation data release
- Data types: teleoperation sequences, egocentric video, physical interaction recordings