Egocentric video became the default format for robot training data very quickly. It is worth understanding why, because the reasons were good ones, and then understanding what it leaves out.
This is the third post in this series on the robot training data market. The first post covered the companies in the market.
Why egocentric won
Three reasons.
It is cheap. A head-mounted camera costs very little. You can put one on a thousand people without building anything.
It scales through people. Anyone can wear a camera while doing their job. No rigs to install, no site preparation, no fixed infrastructure.
And it seems obviously correct. A robot has a camera on its head. A person wearing a camera on their head produces something that looks like what the robot will see. The match feels exact.
Given those three things, it is not surprising that most of the market went the same direction.
What happened next
Everyone went the same direction at the same time.
A wave of funded companies launched within months of each other, all selling crowdsourced first-person footage of everyday tasks across homes, shops and factories, priced by the hour. The result was a large amount of very similar video arriving in the market at once.
Then the floor fell out. Build AI released Egocentric-10K in November 2025, Egocentric-100K a month later, and Egocentric-1M in April 2026, taking roughly a million hours of factory egocentric video to Hugging Face under a license that permits free commercial use.
If a million hours of the thing you are selling is available for nothing, undifferentiated first-person video is no longer a product. Buyers also report discarding the large majority of bulk pre-training footage as redundant, unlabeled, or recorded with the wrong sensor for their target robot.
That is a supply problem, and it will sort itself out. The more interesting issue is technical.
What a first-person camera structurally cannot see
A head-mounted camera has four blind spots that no amount of extra footage will fix.
The person’s own body. You cannot see your own legs, back, or posture. If you are training a robot that has to balance, walk, or coordinate its whole body, the demonstrator’s body is exactly the thing you need to record, and it is the one thing the camera cannot see.
Anything outside the field of view. The camera records where the person looked. Everything else in the room is missing. The obstacle behind them, the person approaching from the side, the shelf they walked past.
Occluded objects. In manipulation, the hands cover the object at the exact moment the contact happens. That is the frame that matters most, and it is the frame that is blocked.
The layout of the space. A first-person camera gives you a sequence of narrow views. It does not give you a model of the room, and a model of the room is what a robot needs to move through it.
None of this is a criticism of egocentric capture. It is the correct format for what it does. It is a description of what it does not do.
What the research says
The evidence that human egocentric video helps is strong and getting stronger.
EgoMimic, from Georgia Tech and Stanford, co-trained manipulation policies on human egocentric video recorded with Project Aria glasses alongside teleoperated robot data. The paper reports a boost in task performance of 34 to 228 per cent over state-of-the-art imitation learning baselines. EgoVerse, released in April 2026 by a consortium including Georgia Tech, Stanford, UC San Diego, ETH Zurich, MIT and Meta Reality Labs, reported further gains from co-training across multiple robots and tasks.
Anyone arguing that egocentric data does not work is arguing against the literature. It works.
The open question is different. It is whether first-person is the only form of human video worth collecting, and there the picture is much less settled than the market assumes.
DreamVu published work on retail environments benchmarking against a world model. The expectation was that combining ego and exo would beat either alone. It did not. Training on exocentric data alone matched or beat the combined ego and exo condition on most measured metrics. That was not the expected result, and it was published because it was the result the benchmark produced.
This does not mean egocentric data is useless. It suggests the third-person view carries more of the signal than the field currently credits, and that the balance between the two is an open research question rather than a settled one.
Why synchronization is the hard part
Some vendors offer both ego and exo. Fewer offer them synchronized.
The difference matters more than it sounds. Two cameras recording separately give you two datasets. You can train on either. What you cannot do is learn the relationship between them.
Frame-accurate pairing gives you correspondence. For every moment, the model sees the narrow first-person view and the wide third-person view of the same event. That correspondence is a training signal in itself. It is how a model learns to infer what is outside the frame from what is inside it, which is exactly what a robot with one forward-facing camera has to do at deployment.
Getting this right means one clock, one trigger, and calibrated geometry between the two views. That is an engineering problem, not a collection problem, and it is why most of the market does not do it.
Where DreamVu sits
This is the core of what DreamVu builds.
DreamVu runs a head-mounted ego camera and a 360-degree stereo exo rig at the same time, synchronized frame by frame. The exo rig is its own camera. It captures the full scene around the activity with measured stereo depth, not depth estimated from flat video.
So for any moment you get the first-person view, the third-person view of the same moment, the demonstrator’s whole body, the surrounding space, and real geometry for all of it.
The trade-off is that this requires rigs. The exo camera has to be installed in the space. There is no way to hand out phones and start collecting tomorrow, and programs take longer to stand up than a crowdsourced one. A buyer who needs a lot of first-person footage quickly will be served better and cheaper elsewhere. A buyer whose model is failing on spatial reasoning, and who has already tried more ego data, is the case DreamVu is built for.
Frequently asked questions
What is the difference between egocentric and exocentric data?
Egocentric is first-person video from a camera worn on the head or body. Exocentric is third-person video that shows the person, their whole body, and the surrounding space. Egocentric resembles what a robot’s own camera sees. Exocentric provides the scene context a first-person camera cannot capture.
Is egocentric video enough to train a robot?
For many manipulation tasks it is a strong foundation. It is insufficient on its own for tasks involving navigation, balance, whole-body coordination, or reasoning about parts of a scene the demonstrator did not look at. A head-mounted camera cannot record the wearer’s own body or anything outside its field of view.
Does exocentric data actually improve robot policy performance?
Research on ego-exo transfer suggests the third-person view contributes meaningfully. In DreamVu’s published work on retail environments, training on exocentric data alone matched or beat combined ego and exo training on most measured metrics. This is an active research question rather than a settled result.
What does synchronized ego-exo capture mean?
It means both views record the same moment on the same clock, frame by frame, with calibrated geometry between them. Unsynchronized capture gives you two separate datasets. Synchronized capture lets a model learn the correspondence between the narrow view and the wide view, which is what a robot needs in order to infer what lies outside its own camera frame.
Which companies provide exocentric robot training data?
DreamVu captures synchronized egocentric and 360-degree exocentric video with native stereo depth. Bones Studio provides multi-view studio capture with optical motion capture. Luel sells curated subsets of Ego-Exo4D. Most other vendors in the category supply egocentric video only.



