日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38046

EgoAlign: 人間とヒューマノイドのギャップを埋める長距離移動操作

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点の人間デモを身体スケールと制御応答の違いを補正してヒューマノイドの全身制御用データに変換し、VLAモデルのファインチューニングと実機ゼロショット展開を実現した。

詳しい要約

1. どんなもの?

- どんなもの? - EgoAlignは、egocentric human demonstrationsをhumanoidの訓練監督に変換するデータ構築フレームワーク。 - 身体スケールやコントローラ応答の違い、robot statesの欠如を克服。 - 物理ロボットのデモンストレーション収集なしで、汎用連続全身コントローラと互換性のあるaction and state supervisionを生成。 - 視覚誘導周期ステッピングのlocomotion referencesを保持しつつ、上半身相互作用ジオメトリをscale alignmentとcontroller-in-the-loop refinementで適応。 - 最終的なcausal replayでrobot statesとmotion-token labelsを再構築し、human observationsと共に訓練。

2. 先行研究と比べてどこがすごい?

- 先行研究と比べてどこがすごい? - 従来は身体スケールやコントローラ応答の違い、robot statesの欠如により、egocentric human demonstrationsのhumanoid訓練への利用が限定的だった。 - EgoAlignは物理ロボットデモなしで、これらのデモをaction and state supervisionに変換可能。 - 提案手法により、シミュレーションでのhand alignmentと物理pickup successがkinematic alignmentのみより改善。 - 人間による収集はteleoperationと比べ現場取得時間を削減。

3. 技術・手法の肝は?

- 技術や手法の肝はどこ? - ターゲットロボットモデルとシミュレータを使用し、実行フィードバックを通じてデモ収集をガイド。 - locomotion referencesを保持し、上半身相互作用ジオメトリをscale alignmentとcontroller-in-the-loop refinementで適応。 - 最終的なcausal replayで対応するrobot statesとmotion-token labelsを再構築。 - 人間の観察と共に訓練するための監督を生成。

4. どうやって有効だと検証した?

- どうやって有効だと検証した? - 適応された人間デモのみでvision-language-action modelをファインチューニング。 - 物理humanoidにゼロショット展開。 - 結果として、長距離物体移動、未見目標位置へのナビゲーション、独立評価の足相互作用を実行。 - refinementがシミュレーションでのhand alignmentと物理pickup successをkinematic alignmentのみより改善。 - 人間収集がteleoperationと比べ現場取得時間を削減。

5. 議論はある?

- 議論はある? - 要旨からは不明。

6. 次に読むべき論文は?

- 次に読むべき論文は? - 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、vision-language-action model、teleoperation、kinematic alignment、controller-in-the-loop refinement、causal replayが挙げられる。 - 同分野の定番として、humanoid loco-manipulation、egocentric demonstration learning、whole-body controlが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen, Aimin Hao, Yisheng He

分類: cs.RO

原文アブストラクト

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/

関連論文

PR本紙発行元 EmplifAI