← Blog

Manipulation Is Not the Whole Robot

Manipulation Is Not the Whole Robot

If you are buying robot training data, there is a gap in the market that is easy to miss until it costs you a quarter.

Most of the data available is about hands.

This is the fourth post in this series on the robot training data market. The first post covered the companies in the market.

The scope ladder

Robot capability is often discussed as one thing. It is at least five, and the data requirements differ at every step.

Fine manipulation. Fingers, grasp, contact forces, in-hand adjustment.

Bimanual coordination. Two arms working together on one object. Holding while cutting. Opening while pouring.

Whole-body coordination. Using the torso, hips and legs as part of the task. Lifting from the floor. Reaching above the head. Bracing against something.

Locomotion. Walking, balance, stairs, uneven ground, recovery from a stumble.

Mobile manipulation. All of the above at once. Walk into a room, find the object, approach it, position the body, pick it up, carry it somewhere else.

A robot arm bolted to a table needs the first two. A humanoid or a mobile manipulator needs all five.

Where the market clustered

Look at what the well-funded companies actually say about themselves.

Config describes itself as building data infrastructure for general-purpose bimanual robotics. That is on its own website. Not a criticism, a description.

XDOF launched with ABC-130K, which it describes as the largest open-source bimanual robot manipulation dataset. Its capture stack is built around teleoperation rigs and hand tracking.

This concentration is rational. Manipulation is the visible bottleneck. It demos well. And it is capturable, which matters more than anything else. You can put a fixed rig on a table and collect all day. You cannot easily follow someone walking around a building with the same equipment.

So the market did the sensible thing and solved the capturable problem first.

What breaks when a mobile robot trains on tabletop data

The failure mode is specific, and it does not look like a data problem at first.

The robot picks up the object correctly on the bench. Then you put it in a room, and it cannot get to the object. It approaches from a bad angle. It positions its body so the arm cannot reach. It stops when a person walks past because it has never seen a person walk past. It reaches for something at floor level and loses balance because nothing in its training involved a torso.

None of these is a manipulation failure. The manipulation is fine. Everything before the manipulation is missing.

Tabletop data contains no gait, no balance, no navigation, no self-model of a body in space, and no experience of approaching a task from across a room.

Bones Studio and the whole-body case

The clearest evidence that locomotion data is a separate discipline is what it takes to produce it.

Bones Studio came from entertainment motion capture. It runs optical motion capture at 120 frames per second with sub-millimetre accuracy, recording 3D motion, multi-view video, audio and 3D scene reconstruction together in a single take.

That data was used to build SONIC, a whole-body control model for humanoid robots. Bones then released BONES-SEED, over 142,000 annotated motion sequences covering locomotion, everyday activities, object interactions and complex whole-body behaviour, in NVIDIA SOMA and Unitree G1 formats.

Note what was required. A motion capture stage. Optical marker systems. Hundreds of performers. This is not something a head-mounted camera produces as a by-product.

For locomotion policy or motion generation, Bones is the strongest source available.

The cell nobody is filling

Bones solves whole-body motion at very high precision. It does so on a stage, with performers, doing directed motions.

That is the right trade for motion fidelity. It is the wrong trade for environment realism.

A performer walking across a capture volume is not a warehouse worker walking down an aisle between stacked pallets while a forklift reverses behind them. The motion is clean, the space is empty, and the task is performed rather than done.

So the market has two things. Very precise whole-body motion in a studio. Very cheap first-person hand footage in the field.

What it mostly does not have is whole-body activity, in real working environments, with real geometry.

Where DreamVu sits

DreamVu produces whole-body data as a consequence of its capture method, not as a separate program.

The exo rig is a 360-degree stereo camera placed in the working environment. It records the person’s entire body, the space around them, and the objects in it, with measured depth. When someone walks across a shop floor, bends to a low shelf, carries a box across a room and sets it down, all of that is in the recording. So is the room they did it in.

That means the data covers navigation and whole-body movement alongside manipulation, in the actual environments where the robot will work.

The trade-off is precision at the fine end. DreamVu does not deliver finger-level dexterity at motion capture accuracy. Bones does that better and it is not close. For dexterous in-hand manipulation, Bones is the correct vendor. For a robot that cannot get to the object in the first place, DreamVu is.

Frequently asked questions

What is the difference between manipulation data and whole-body data?

Manipulation data covers hands, grasping and object contact, usually recorded at a fixed workstation. Whole-body data covers the torso, hips and legs, including walking, balance, bending and reaching, and usually requires either motion capture or a third-person camera view that can see the entire body.

Why is most robot training data focused on manipulation?

Because it is the easiest to capture. A fixed rig at a table can record all day. Manipulation is also the most visible bottleneck in robot capability and it demonstrates well. Following a person moving through a building requires different equipment entirely.

What data do humanoid robots need that manipulation datasets do not provide?

Locomotion, balance, whole-body coordination, and navigation. A humanoid has to walk to the task, position its body correctly, remain stable while reaching, and move safely around people. None of that is present in tabletop manipulation recordings.

Which companies provide whole-body and locomotion data for robots?

Bones Studio is the leading specialist, using optical motion capture, and released BONES-SEED with over 142,000 annotated motion sequences covering locomotion and whole-body behavior. DreamVu captures whole-body activity in real working environments using 360-degree exocentric stereo. Mecka AI records walking gait alongside hand dexterity using body-worn sensors.

Is motion capture data better than video for whole-body training?

For motion accuracy, yes. Optical motion capture is measured to sub-millimeter precision and is more accurate than anything estimated from video. The trade-off is environment. Motion capture happens on a stage with performers, so it does not contain the clutter, obstacles and unpredictability of a real workplace.

Tell us what your model needs to learn.

Capture programs, research collaboration, and dataset partnerships.

Talk to us