日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
全身移動操作arXiv:2609.19340

ViLoMan: ヒューマノイドロボットの視覚・固有感覚による全身移動操作スキルの学習

ViLoMan: Learning Visual-Proprioceptive Whole-Body Loco-Manipulation Skills for Humanoid Robots

シェア:XThreadsFacebookLINEはてブBluesky

人間の部分的な動作デモを物理的に実行可能なロボット軌道に変換し、教師-生徒蒸留で視覚と固有感覚から全身関節動作を直接出力する統合方策を学習。扉閉めタスクでシミュレーションから実機へ転移し、多様な条件で汎化を実証した。

詳しい要約

1. どんなもの?

- 本論文は、ヒューマノイドロボットの自律的なloco-manipulation(移動と操作の統合)のためのスケーラブルなフレームワークViLoManを提案する。 - 部分的で運動学的な人間-物体インタラクションのデモンストレーションを、物理的に実行可能な完全なロボット軌道に変換する。 - 変換された軌道をteacher-student distillationフレームワークで利用し、egocentric depth観測とproprioceptive測定から直接関節レベルの全身行動にマッピングする統一ポリシーを学習する。 - 展開時には参照動作や中間コマンドを必要としない。 - シミュレーションと実世界の多様なドア構成とロボット初期条件でドア閉めタスクを評価する。

2. 先行研究と比べてどこがすごい?

- 従来の研究と比較して、多様で物理的に実行可能なロボット-物体インタラクションデータの不足と、オンボード観測から統一的な全身制御を直接学習することの難しさを克服する。 - ViLoManは、部分的な運動学的デモンストレーションを完全な物理的軌道に変換し、teacher-student distillationを活用することで、参照動作や中間コマンドなしで動作する統一ポリシーを実現する。 - 単一のポリシーで多様なタスク変動にロバストに一般化し、シミュレーションから実世界への効果的な転移を可能にする点が優れている。

3. 技術・手法の肝は?

- 部分的で運動学的な人間-物体インタラクションのデモンストレーションを、物理的に実行可能な完全なロボット軌道に変換する。 - 変換された軌道をteacher-student distillationフレームワーク内で利用し、egocentric depth観測とproprioceptive測定を入力として、関節レベルの全身行動を出力する統一ポリシーを学習する。 - 展開時には参照動作や中間コマンドを必要とせず、オンボードのdepth sensingとproprioceptionのみで動作する。

4. どうやって有効だと検証した?

- シミュレーションと実世界の両方で、多様なドア構成とロボット初期条件のドア閉めタスクを評価した。 - 実験結果は、単一のポリシーがUnitree G1ヒューマノイドにオンボードのdepth sensingとproprioceptionのみを使用して完全なタスクを完了させ、タスク変動にロバストに一般化し、シミュレーションから実世界へ効果的に転移することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法として、teacher-student distillation、loco-manipulation、whole-body control、sim-to-real transferなどの同分野の定番が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zejie Tian, Ruibing Hou, Bingpeng Ma, Börje F. Karlsson, Shiguang Shan

分類: cs.RO

原文アブストラクト

Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.

関連論文

PR本紙発行元 EmplifAI