日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08726

EgoLAP: 言語-行動推論による一人称視点の人間データからの学習

EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

人間の一人称視点の軌道を、言語による行動連鎖推論を介してロボット制御に転移させるVLA事前学習フレームワークを提案し、実世界タスクで大幅な性能向上を達成した。

詳しい要約

1. どんなもの?

EgoLAPは、egocentric human dataを活用してrobot learningをスケールさせるVLA pre-training framework。人間とロボットのtrajectoryから、共有されたlanguage-based action chain-of-thoughtを通じて共同学習する。motion intentを構造化・時間抽象化されたlanguage actionとして表現し、scene geometry・physics・object affordancesに基づくmotion-level reasoningと組み合わせる。

2. 先行研究と比べてどこがすごい?

従来はembodiment gapのため生の人間trajectoryは制御のsupervisory targetとして不適だった。EgoLAPは低レベルactionはembodiment固有でもmotion intentは人間とロボット間で転移可能なtask-relevant structureを捉える点を活用。代替action representationより人間経験をロボット制御に効果的に転移し、実世界タスク進捗80.1%平均、代替比2.3倍の性能向上を達成。

3. 技術・手法の肝は?

VLA pre-training frameworkで、人間とロボットのtrajectoryを共有のlanguage-based action chain-of-thoughtで共同学習。motion intentを構造化・時間抽象化されたlanguage actionとして表現し、scene geometry・physics・object affordancesに基づくmotion-level reasoningとペアにする。

4. どうやって有効だと検証した?

広範な実世界およびシミュレーション実験を実施。EgoLAPは代替action representationより人間経験をロボット制御に効果的に転移し、実世界タスク進捗80.1%平均、代替比2.3倍の性能向上。motion-level reasoningはsubtask・object-box・visual-trace reasoningを組み合わせたcomposite reasoning formatより優れる。

5. 議論はある?

motion-level reasoningがcomposite reasoning format(subtask・object-box・visual-trace reasoningの組み合わせ)を上回ることが示された。その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されているのは代替action representationsとcomposite reasoning format(subtask, object-box, visual-trace reasoning)。関連手法としてVLA pre-trainingやegocentric human dataからの学習が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lihan Zha, Shresth Grover, Tenny Yin, Samuel M. Bateman, Hengkai Pan, Mengchao Zhang, Aykut Onol, Allen Z. Ren, Dhruv Shah, Anirudha Majumdar

分類: cs.RO, cs.AI

原文アブストラクト

Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.

関連論文

PR本紙発行元 EmplifAI