日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11283

Being-M0.7: ヒューマノイドロボットのための潜在世界行動モデル

Being-M0.7: A Latent World-Action Model for Humanoid Robots

シェア:XThreadsFacebookLINEはてブBluesky

人間の動画・モーションの大規模データから視覚運動の事前知識を学習し、ヒューマノイドの全身移動・操作を実現する世界行動モデルを提案。シミュレーションと実機Unitree G1で高い成功率を示した。

詳しい要約

1. どんなもの?

- ヒューマノイドロボットの loco-manipulation のための latent world-action model である Being-M0.7 を提案。 - 人間の video と motion の混合モダリティデータから視覚運動 prior を学習し、humanoid 制御に転移。 - pre-training、robot mid-training、action post-training の3段階で構成。 - 10,000時間以上の人間中心データを収集し、video-only、motion-only、paired video-motion を統合。 - 将来の latent visual state と motion の同時予測により、視覚表現に将来の運動情報をエンコードさせる。

2. 先行研究と比べてどこがすごい?

- 従来はロボット実演データが乏しく、人間の video や motion データは video か motion の片方のみが多い。 - また人間の motion はそのままロボットの実行可能な action にはならないという課題があった。 - 本研究は混合モダリティの人間データから視覚運動 prior を学習し、humanoid 制御に転移する点が新しい。 - 比較ベースラインの中で SIMPLE において最高の aggregate success rate を達成。 - 実世界の Unitree G1 loco-manipulation タスクでは最強ベースラインと同等の性能を示した。

3. 技術・手法の肝は?

- 10,000時間以上の人間中心データから video-only、motion-only、paired video-motion を統合したコーパスを構築。 - pre-training では将来の latent visual state と motion を同時予測し、視覚表現に将来の運動情報をエンコード。 - robot mid-training で粗い prior をロボットの視点と身体動態に適応。 - action post-training では action expert が、凍結・適応済み prior からの視覚予測表現と現在の画像・proprioception を gated cross-attention で統合。 - 予測コンテキストを実行可能な全身コマンドに接地させる。

4. どうやって有効だと検証した?

- SIMPLE ベンチマークで比較ベースライン中最高の aggregate success rate を達成。 - 実世界の Unitree G1 loco-manipulation タスクで最強ベースラインと同等の性能を確認。 - これにより、提案手法の有効性をシミュレーションと実機の両面で検証。

5. 議論はある?

- 要旨からは、限界や議論の詳細は不明。 - 混合モダリティデータの活用や転移学習の有効性は示唆されるが、失敗事例や計算コストなどは言及されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法として、human video や motion からのロボット学習、latent world model、humanoid loco-manipulation の定番研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junpeng Yue, Boyuan Li, Yuxuan Wang, Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang, Jing Zhang, Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.

関連論文

PR本紙発行元 EmplifAI