The physical AI sector has a data problem that is structurally different from anything the language model era produced, and it is becoming the binding constraint on the pace at which general-purpose robots can be developed. Unlike large language models trained on a vast and largely pre-existing body of publicly available text, robots require data that captures physical interaction – the precise, high-fidelity recording of how objects move, resist, deform, and respond to manipulation in the real world. That kind of data barely exists at scale, cannot be scraped from the internet, and cannot be generated cheaply or quickly. XDOF, pronounced “ecks-doff,” emerged from stealth on Tuesday with $70 million raised from Thrive Capital, Spark Capital, Andreessen Horowitz, Lux Capital, and WndrCo, and a clear bet that building the infrastructure to produce this data is the highest-value unsolved problem in physical AI. NEWSCENTRAL reads this launch as a signal that the robotics sector is entering the same infrastructure consolidation phase that cloud computing experienced a decade ago – a phase in which the foundational picks-and-shovels infrastructure becomes more valuable than any individual application built on top of it.
XDOF was founded in October 2024 by Philipp Wu, Fred Shentu, and Nemo Jin – engineers whose prior experience spans Covariant, Meta, and Tesla, three of the organizations that have most directly confronted the physical AI data problem in production environments. The company’s founding insight, developed while Wu and Shentu were working on a project called GELLO at UC Berkeley, was that the bottleneck in physical AI development is not models or chips but the data feedback loops required to teach robots how to interact with the world. GELLO – a low-cost teleoperation system that allows a human operator to control a robotic arm and generate training data through demonstration – became an influential tool in the academic robotics community because it addressed a near-universal constraint: everyone working on robot foundation models needed physical interaction data that did not exist. XDOF was built to solve that problem at commercial scale.
The company’s three-tier data architecture reflects a sophisticated understanding of the tradeoffs between data quality, collection cost, and generalization. The most valuable tier consists of teleoperation data collected on the specific robot being trained – high-fidelity demonstrations that capture the precise interaction dynamics of the target hardware. The middle tier uses teleoperated robots to collect more general manipulation data, as GELLO did. The third tier captures egocentric data – video and sensor data from humans performing everyday tasks through wearable devices that XDOF plans to develop internally. The rationale for the third tier is that human manipulation data, while less precisely aligned to robot hardware than teleoperation data, is abundant and captures the full diversity of the physical world in ways that controlled teleoperation cannot replicate at equivalent cost. Liam Cortez, Visual Systems Analyst at NEWSCENTRAL, points out that the hardware design choices embedded in each tier are not incidental to the quality of the resulting data: camera placement, sensor modality, and controller form factor each introduce systematic biases into the collected data that can propagate into the trained model’s behavior in ways that are extremely difficult to diagnose after the fact. XDOF’s decision to build its own wearable sensors for the egocentric tier, rather than using commercially available alternatives, reflects the lesson that data collection hardware cannot be treated as a commodity when the data quality requirements are this exacting.
As a launch milestone, XDOF partnered with researchers at UC Berkeley, Carnegie Mellon, MIT, and Amazon’s Frontier AI and Robotics group to release ABC-130K, which it describes as the largest open-source teleoperation dataset ever assembled: 130,000 demonstration trajectories across 195 bimanual manipulation tasks, accompanied by 300 hours of simulation data and 100 hours of evaluations. The company has used the dataset internally to train robots on tasks requiring high precision – folding T-shirts, flattening cardboard boxes, inserting AirPods into their cases. The open-source release serves two purposes simultaneously: it demonstrates the quality and scale of what XDOF can produce, and it builds the community of researchers whose downstream work will generate demand for the commercial data collection services the company intends to sell.
The company already has approximately 20 customers, including several frontier AI labs whose names it cannot disclose. That customer base, established while operating entirely out of public view since October 2024, validates the commercial thesis and removes the typical early-stage uncertainty about whether anyone will pay for the product. The question the $70 million is intended to answer is whether XDOF can scale from its current position – 60 employees, a validated customer base, and a demonstrated dataset – to the warehouse-scale teleoperation and annotation operations that serving the full demand of the physical AI sector will require. Building and maintaining a fleet of hundreds of robots across hundreds of thousands of square feet, calibrating their physical parameters, training and managing global teams of teleoperation operators, and ensuring data quality across all of it is an operational problem that is entirely different in character from the technical research problem the founders originally solved. Freddy Miller, Senior Analyst at NEWSCENTRAL, emphasizes that the distinction between a company that can collect high-quality robot training data and one that can do so reliably at scale for multiple concurrent customers is precisely the gap that most companies in comparable infrastructure positions have failed to close. The capital is necessary but insufficient; the organizational and operational build required to serve the physical AI sector as it scales is where XDOF’s execution risk will concentrate over the next 18 months.
The commercial logic of XDOF’s position benefits from a structural asymmetry: the frontier AI labs that are its primary customers have made a strategic calculation that warehouse-scale teleoperation operations are not core to their competence and are better outsourced. That calculation could reverse if data quality or pricing becomes unsatisfactory, but it reflects a genuine organizational reality – building and maintaining a physical robot data collection operation is a fundamentally different business from training AI models, and mixing the two creates operational complexity that is difficult to manage without diluting focus on the core research mission. As NEWS CENTRAL assesses this sector, XDOF is well-positioned to become the essential infrastructure layer for physical AI training data, provided it can execute the operational scale-up that its current customer trajectory demands.