日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00438

自己中心的全身体人間データの事前学習による汎用人型ロコマニピュレーションモデル

Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining

シェア:XThreadsFacebookLINEはてブBluesky

ウェアラブルシステムで収集した500時間の自己中心的人間行動データセットHumanVerse-500を構築し、それを用いた3段階学習で人型ロボットの全身視覚言語行動ポリシーλ0を開発した。

詳しい要約

1. どんなもの?

- ヒューマノイドの全身 loco-manipulation を、egocentric な人間の動画からスケーラブルに学習する試み。 - 500時間の人間全身行動データセット HumanVerse-500 を構築。 - 軽量ウェアラブルシステムで egocentric video と身体・手の動きを同期収集。 - これを用いて全身 humanoid vision-language-action policy である λ0 を開発。 - 3段階学習で人間経験をロボットに転移する。

2. 先行研究と比べてどこがすごい?

- 従来の egocentric 動画監督は全身運動と hand-object interaction の協調を十分にカバーしていなかった。 - humanoid teleoperation による監督は高コストでスケールしにくい。 - 本研究は人間経験をスケーラブルに活用し、全身 loco-manipulation 学習を可能にする。 - HumanVerse-500 は open-world 環境での多様な人間 loco-manipulation 行動を500時間分含む。 - λ0 は SIMPLE と4つの実世界タスクで state-of-the-art 性能を達成。

3. 技術・手法の肝は?

- 3段階学習を採用。 - 第1段階:多様な egocentric データセットから interaction を学習。 - 第2段階:HumanVerse-500 を用いて身体と手の動きを協調。 - 第3段階:下流タスクとロボット embodiment に適応。 - 人間経験転移のための共有表現空間を学習。 - 人間とロボットの状態・行動の差異はドメイン固有インターフェースで処理。

4. どうやって有効だと検証した?

- SIMPLE と4つの実世界 loco-manipulation タスクで評価。 - state-of-the-art 性能を達成。 - スケーリング挙動、汎化、学習段階ごとの寄与を分析。 - 人間データが下流の全身 humanoid 制御をどう支えるか検討。

5. 議論はある?

- 人間データが下流の全身 humanoid 制御をどう支援するかを分析。 - スケーリング挙動と汎化を検証。 - 各学習段階の寄与を評価。 - コード、モデル、データを公開予定。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として humanoid teleoperation、egocentric video からの学習、vision-language-action policy が挙げられる。 - 同分野の定番として humanoid loco-manipulation、imitation learning、vision-language-action models が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chongyang Xu, Zhao Wu, Jin Chen, Yiming Jiang, Jinhui Ye, Yuming Jiang, Shifeng Zhang, Ziliang Feng, Mu Xu, Yilun Chen, Li Lu, Steven C. H. Hoi

分類: cs.RO

原文アブストラクト

Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop $λ_0$, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, $λ_0$ learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate $λ_0$ on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.

関連論文

PR本紙発行元 EmplifAI