Large models can “read the entire internet,” but robots must practice repeatedly in the physical world. Picking up a cup, folding clothes, inserting and connecting parts — these actions, which seem second nature to humans, require a robot not only to record the visuals of each step, but also to simultaneously record joint positions, action commands, force sense, tactile sense, and task results. Data collection has thus become the key infrastructure connecting the robot body, sensors, models, and real-world scenarios, and is attracting growing attention across the industry.


The Scale of the Data-Collection Market: Why Does Data Matter?

At present, when discussing “how big is China’s embodied-intelligence data-collection market,” the greatest taboo is to directly cite the scale of humanoid robot whole-machine units. Although some third-party research says the global embodied-intelligence dataset market will grow from USD 737 million in 2024 to USD 7.014 billion in 2031, a compound growth rate of 38.2%, with China accounting for about half of the global total by then. The industry currently still mixes together metrics such as number of data entries, number of trajectories, effective hours, equipment orders, and data-service revenue, and public datasets can also be sold repeatedly, so the measurement caliber has not yet converged.

On the demand side, two clear signals have already appeared:

The fundamental force driving growth is that data scale and quality can already directly change model performance. Embodied models cannot, like language models, directly obtain massive, ready-made internet training corpora. For the very same “pick up a cup” task, the robot must both understand the environment and the command, and also learn the movement trajectory, the grasping force, and the recovery strategy after a failure; and if the robot’s configuration, camera position, or end tool changes, the original data may no longer be directly reusable. By means of large-scale, high-quality data, NVIDIA’s EgoScale used 20,854 hours of first-person human video for pre-training, raising the average success rate by 54% over a no-pre-training baseline on a 22-degree-of-freedom dexterous-hand task.

The second is that domestic public capacity is forming at an accelerated pace; accompanying the construction of training grounds and the advancement of customized projects, data-collection equipment sales and the improvement of the data-collection industry chain have become a short-term driver. The Shanghai Zhangjiang training ground deployed more than a hundred heterogeneous robots in its first phase; the Beijing Shijingshan data-training center is equipped with 100 robots and 100 sets of data-collection equipment, and plans an annual output of over one million entries of multimodal real-world data.


Data-Collection Technology Routes: Data for Different Levels

Closest to the robot’s real execution state is real-machine teleoperation. Operators control the robot through master-slave robotic arms, exoskeletons, VR handles, or data gloves, simultaneously collecting RGB-D, joint angles, end poses, torque, and touch. At present, industry solutions are evolving toward backpack-based, modular, and low-cost designs. This type of data is the most consistent with the target body, with the highest precision and information density; the price is expensive equipment, slow scene resetting, and weak cross-body transfer. Referring to the data-collection case of a certain domestic leading humanoid robot vendor, the cost is equipment investment of more than RMB 200,000 per set (approx. USD 29,500), with data at about RMB 500–1,000 per effective hour (approx. USD 74–147).

If teleoperation collects “how the robot does it,” the Ego first-person route records “how the human does it.” The collector wears a head-mounted or wrist-mounted RGB/RGB-D camera, and can additionally layer on an IMU, data gloves, and audio, operating naturally in real environments such as homes and factories. Ego does not require a robot body, and has the strongest scene diversity and scaling capability, with a combined cost about 1/5 that of teleoperation; but the output is mainly human video, hand poses, and trajectories, which require action retargeting before they can be converted into robot commands, and are also susceptible to factors such as occlusion, drift, clock desynchronization, and the absence of force and tactile sensing.

UMI lies between the two: it neither directly controls a real machine, nor merely films the human hand, but instead has the collector hold a “robot gripper” to complete the demonstration. A camera, IMU, and open-close sensor are built into the handheld gripper, recording the end six-dimensional pose, gripper opening and closing, and operation video through visual SLAM fused with inertial data. It is closer to the robot’s action space than Ego, with less post-processing, and is more portable and cheaper than real-machine teleoperation, with equipment investment at the ten-thousand-yuan level (approx. USD 1,500 level). However, so-called “universality” has a clear precondition: if the robot’s end gripper, camera model, and calibration parameters are inconsistent with the collection equipment, transfer effectiveness will decline, and it is also hard to cover the fine operations of a dexterous hand.

Lower in cost and faster to scale are simulation and synthetic data, which can be further subdivided into physics-engine simulation, Real-to-Sim scene reconstruction and trajectory/asset synthesis, and generative world models. The advantage is that dangerous, rare, and extreme working conditions can be manufactured in batches; the shortcoming is that physical details such as friction, flexible materials, and contact force are hard to fully reproduce, and a Sim-to-Real gap always exists. Caitong Securities, citing data from Songying Technology, said that the cost of a single entry of real-machine data is about RMB 3–5 (approx. USD 0.44–0.74), while simulation is about RMB 0.2–0.3 (approx. USD 0.03–0.04).


From Practical Cases: Where Does the Waste in “Effective Data Cost” Occur?

What data-collection projects most easily underestimate is often not the equipment price, but the effective rate. According to BCC Research, one 8-hour shift of whole-machine teleoperation requires deployment, calibration, troubleshooting, and resetting, and after cleaning there may remain only about 1 hour of effective data (by contrast, UMI can reach 4–5 hours, with a combined daily cost per workstation of about RMB 300–350 [approx. USD 44–52]). This reminds relevant data purchasers: one should not look only at equipment price or time on duty, but rather calculate unit cost by “effective trajectories that pass quality inspection,” comprehensively considering the pass rate, output efficiency, or rework rules.

Another type of loss is hidden in the details of equipment and environment. An expert interviewed by BCC explained that head-mounted devices may, due to structural loosening, heat-induced deformation, optical distortion, electromagnetic interference, or mid-way power loss, cause misalignment between the visual and IMU coordinates; scenarios such as deep cabinets or backlighting can also cause both hands to leave the field of view. Each company is exploring corresponding improvement methods, and some enterprises will even write about 10% of controllable failures into the SOP, using them to train anomaly recognition and recovery. The insight here is not to pursue “zero failure,” but to distinguish between failures that can teach the model to recover and bad data that only creates noise.

After entering industrial scenarios, long-tail working conditions push the cost up by another level. A certain enterprise supply-chain case in BCC Research shows that the purchase price and effective rate differ markedly across simple quality inspection, 3C loading and unloading, and new-energy battery processes; low-frequency working conditions such as reflection, slipping, cable tangling, and battery bumping are hard to collect repeatedly on a real machine, and the effective rate for high-risk processes may be only about 25%. At this point, the calibration between simulation and real data becomes an effective solution, and is also where the capabilities of different data-collection suppliers / embodied-intelligence body vendors are distinguished.

Looking to the future, China’s embodied-intelligence data collection is shifting from “building factories and piling up hours” toward “operating data by outcome.” The key to industry competition is no longer merely how many collectors and how much equipment one has, but who can turn task design, multi-sensor synchronization, quality assessment, cross-body conversion, and model back-testing into a closed loop. In measuring a data-collection company, one should also return to three questions: how high the effective rate is, whether it can transfer, and how much it can raise the robot’s success rate.

[Disclaimer]: The above content reflects analysis of publicly available information, expert insights, and BCC research. It does not constitute investment advice. BCC is not responsible for any losses resulting from reliance on the views expressed herein. Investors should exercise caution.