Beijing Humanoid Data Hub Surpasses 20 Million Downloads; Director Xia Hualin: The Ultimate Battleground in Embodied AI Will Be the Data Flywheel

Deep News
Yesterday

On the evening of September 3rd, the Embodied AI Data and Training Base of the Beijing Humanoid Robot Innovation Center was officially opened to the public. During a group interview, the base's director, Xia Hualin, stated that the most critical competitive factor in the future of embodied intelligence will not be hardware or algorithms, but rather the data flywheel.

"I come from an autonomous driving background. The reason autonomous driving was deployed so quickly wasn't the vehicles, nor was it the large model algorithms—those were already mature. What was truly needed was data," Xia Hualin pointed out. "This data flywheel will become the most fiercely contested area in the embodied intelligence field, even more so than hardware or algorithms."

The Embodied AI Data and Training Base of the Beijing Humanoid Robot Innovation Center, established in less than a year, covers a total area of over 6,000 square meters. It encompasses more than 30 typical scenarios across six major categories, including home, retail, office, industrial, pharmaceutical, and eldercare. The facility is equipped with over 150 types of embodied hardware and more than 100 sets of non-embodied data collection devices, with an annual production capacity of 180,000 hours of high-quality data.

In terms of delivery, the base has already supplied over 30,000 hours of high-quality data to leading industry clients. Notably, its open-source dataset, RoboMIND, has seen global downloads double within a single month, surpassing the 20 million milestone and ranking first in the industry.

"Our primary focus at Beijing Humanoid is to first serve our internal Tiangong clients, providing data, platforms, SaaS, and, of course, data solutions. However, we are not limited to internal use; we also provide data support to leading industry clients," Xia Hualin explained. "On the commercialization front, there are several avenues—such as selling data. We will sell our high-quality data, as well as offer our platform's service capabilities and data factory solutions. That said, our current goal is not to profit from this, but rather to expand and strengthen the industry and serve its ecosystem."

Regarding the definition of "high-quality data," Xia Hualin noted that the most valuable data is that which allows models to perform effectively and produce strong results in each specific scenario. He illustrated this with an example: "A robot can wash dishes at house A, but fails at house B. The data from the failure at house B is quickly fed back, trained, and deployed, ultimately enabling the robot to perform well at house B as well. It's like Tesla's shadow mode—the data that returns is the highest-value data."

Currently, the base has established a fully self-developed data infrastructure, integrating the entire pipeline of collection, cleaning, annotation, quality inspection, training, feedback, and delivery. Through automated annotation powered by AI large models, data production efficiency has increased by over 50%, while comprehensive costs have been reduced by more than 30%. "In the first half of this year, the data acceptance pass rate for a leading client delivery exceeded 95%," Xia Hualin stated.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10