After 20,000 trial users, China’s robot-data platform starts taking orders

On 23 September MiFengPai, a crowdsourcing app for robot training data, launched publicly. Three weeks earlier it was in closed test, with open questions about where tasks came from, how devices were managed, whether real sites opened, how data passed check and how pay was settled. The first rules are now visible.

In one month of testing the app drew 20,000 registered users and 13,000 submitted collection tasks, covering 22 scenario classes, over 5,000 tasks and more than 50,000 real environments. Zhang Zhifu, the platform’s business head, calls today’s MiFengPai still a small baby, so rather than say the robot Uber is running, it has finally started taking orders.

Crowdsourcing first needs rules for how work is done. Zhang summed up the first half as build to demand, pushing devices and data out to find need, and the firm is now moving to sell to demand, collecting only what is ordered. Chief executive Yao Maoqing put it plainly, the model is demand-led, taking an order then collecting, not blindly gathering.

MiFengPai app interface for robot data collection tasks
The MiFengPai app listing robot-data collection tasks after public launch. (Source: Gasgoo)

Tasks are split by job, scene and action, across logistics, warehousing, care, dining, retail, car service, health and farming, then down to finer moves. Device rules came from the test. Free loans risked idle gear, so a use fee now filters users. The headset standard fee is 39 yuan a day, with launch discounts and credit-based deposit waivers, and wrist gear and tactile gloves will follow.

On what counts as valid, the platform auto-cuts clips with hands off frame, long pauses or empty repeats, then pays by valid time, after which data passes labelling, path extraction and human checks. Yao says over 95 per cent of data survives front-end checks, a figure worth more than the user count. A China Academy of Telecommunication Research report in August warned that scattered crowdsourcing can drag valid rates down through clumsy moves, wrong task reading and messy scenes, so holding 95 per cent as scale grows is the real test.

Diagram of robot training data flowing from collection to delivery
How raw collection becomes delivered robot-training data on the platform. (Source: Gasgoo)

Demand is shifting. Yao noted everyone built home folding-clothes data, now flooded and unwanted, while harder niche skills like plumbing, car repair, haircuts and nails need building out one by one. When easy home and tabletop scenes are over-collected, need moves to professional, scattered long-tail skills. Zhang splits task grain into five types, from a room-scale operation down to a three-second open-the-fridge move, and likens the work to harvesting wheat, where the sold good is bread, not grain.

The US gives a sharper benchmark. Figure’s Index platform, public since August, spans 108 countries, over 264,000 downloads, more than 44,000 weekly active contributors and 16 million uploaded videos, paying 15 million dollars to creators. By 11 September Figure chief Brett Adcock said weekly active contributors passed 86,000, nearly double. In Helix 2.5 tests on 30 unseen homes, adding Index pretraining lifted zero-shot task success from 9 per cent to 56 per cent, and as Index data doubled the model’s next-action prediction followed a smooth scaling law.

The hard part stays reaching the real world. A food plant runs valuable moves daily, yet that does not mean a platform can walk in. Device safety clearance, firm willingness to admit crowdsourced workers, and logistics for remote sites like coconut picking are three gates. The Scenario Data Alliance, with over 50 members including hotels, travel groups, care homes, hospitals and manufacturers, stocks the ability to enter the real world early.

Robot data scenario alliance members across industries
The scenario alliance spans hotels, care, health and manufacturing sites. (Source: Gasgoo)

Privacy and safety rise with entry to hospitals, plants, hotels and homes. Yao said desensitisation is automatic, with back-office staff not seeing raw personal data first, and Zhang said the platform builds layered protection from user consent through on-device processing, cloud desensitisation and compliance checks.

One hard question stayed open. If a mechanic wearing MiFeng gear completes a teardown, the skill is the mechanic’s, the scene and craft the firm’s, the gear and processing the platform’s, and the data may sell to one or more model firms, who shares the value. MiFeng pays the collector per task once, and Yao noted common data may resell to cut repeat collection cost, but whether the skill source and scene owner get only the first payment is unanswered.

Scaling turns hard on devices, supply chain, operation and money. Yao estimates 10 million hours of data needs tens of thousands of device sets, and 100 million hours perhaps 100,000. At about 4 valid hours per device per day and 250 days a year, 10 million hours is roughly 10,000 sets at full load, rarely met in practice. MiFeng’s fixed-asset spend for 10-million-hour scale already runs to several hundred million yuan.

MiFeng is not alone. JD.com launched a full-chain embodied-data base this year and may mobilise 600,000 people to gather 10 million hours in two years, leaning on its retail, logistics, industrial and health scenes. One day before MiFeng’s launch, Yuanli Infinite’s Galaxy Data committed 500 million yuan to a 10-million-hour plan. The paths differ, JD starts with scenes, MiFeng with a third-party platform, Yuanli with its own embodied brain, but all race to turn the real world into a steady data supply chain.

The contest lands on delivery. Yao said top clients moved from trialling dozens of vendors to consolidated buying or single-vendor awards, with orders now at 50,000 or 100,000 hours a week. At that size, being able to collect is only the first gate. The gap is whether, when demand arrives, you can quickly find people, gear and scenes and deliver on time and on spec.

A month ago the question was who makes 100 million hours of data. Now 20,000 people have started answering who makes it. But taking orders and truly running are not the same. Demand needs steady repeat orders, supply needs to keep finding people, gear and scenes at scale, and the business needs clients still buying and collectors still working once subsidies fade. Only when all three loops close does the robot Uber truly run.

Editor’s note: This is an adapted translation of the original Gasgoo report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.gasgoo.com/apps/50640d4b55d5cba175fb84f15d679f19/robot/news/70473070-2%E4%B8%87%E4%BA%BA%E8%AF%95%E6%B0%B4%E4%B9%8B%E5%90%8E/.

Leave a comment