日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.29310

EgoSpeedUp: 人間の操作テンポをロボット方策に転移する

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

人間の操作デモからタスクに適したフェーズごとの実行テンポを推定し、ロボットの遅いデモを再タイミングして模倣学習させることで、成功率を上げつつ実行時間を短縮するフレームワーク。

詳しい要約

1. どんなもの?

- ロボットの模倣学習ポリシーに、人間の操作テンポを転移する枠組み EgoSpeedUp を提案。 - 遅いロボット実演と同一タスクの人間実演を入力とし、位相ごとのテンポを推定してロボット実演を retiming。 - retiming 後の実演で behavior cloning を行い、実行可能な操作を保ちつつ人間由来のテンポで遂行。 - 実世界2タスクで成功率平均+25pp、成功実行時間36.5%減。

2. 先行研究と比べてどこがすごい?

- 既存の加速手法はロボット側情報や事前定義された tempo factor 集合から加速率を決めており、各操作位相に適した基準が不明瞭。 - EgoSpeedUp は人間実演を時間的 supervision として用い、タスクに適した位相ごとのテンポ参照を人間から得る点が新しい。 - ロボット実演の実行可能な挙動を保持しつつ、人間に基づくテンポを学習できる。

3. 技術・手法の肝は?

- 遅いロボット実演と同一タスクの人間実演を用意。 - 対応する操作位相を align し、複数の人間実演から相対的な実行テンポを推定。 - 得られた位相ごとのテンポでロボット実演を retiming。 - retiming 済み実演に対し標準的な behavior cloning を適用。

4. どうやって有効だと検証した?

- 実世界の操作タスク2つで評価。 - 成功率が平均25 percentage points (pp) 向上。 - 成功実行時間が36.5%短縮。 - これにより人間の操作テンポが、より速く信頼性の高いポリシー学習の有効な時間参照となることを示した。

5. 議論はある?

- 人間の操作テンポがロボットポリシーの時間参照として有効であると主張。 - 既存手法がロボット側情報や事前定義 tempo factor に依存するのに対し、人間実演から位相ごとのテンポを得る利点を提示。 - 限界や失敗事例、一般化可能性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている既存の加速手法(ロボット側情報や事前定義 tempo factor を用いるもの)に関する論文。 - 模倣学習の基盤手法である behavior cloning の代表的論文。 - 人間実演をロボット学習に活用する関連研究(例: 人間の動作を supervision とする手法)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hanbit Oh, Yukiyasu Domae, Takuma Yagi

分類: cs.RO, cs.CV

原文アブストラクト

Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.

関連論文

PR本紙発行元 EmplifAI