日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.37599

BlenDAgger: インタラクティブ模倣学習のためのブレンド共有制御

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

シェア:XThreadsFacebookLINEはてブBluesky

介入時に方策と人間の動作を共有制御でブレンドしてデータ収集するBlenDAggerを提案し、実世界2タスクでHG-DAggerより30ポイント以上高い自律性能を達成した。

詳しい要約

1. どんなもの?

- ロボットの模倣学習ポリシーを訓練するためのデータ収集手法である BlenDAgger を提案。 - 介入時にポリシーとデモンストレータの行動を共有制御でブレンドする。 - これにより、操作ポリシーの自律性能を向上させることを目指す。 - 実世界2タスクとシミュレーション3タスクの計5タスクで検証。

2. 先行研究と比べてどこがすごい?

- 典型的な人間ゲート型修正アプローチ (HG-DAgger) と比較して、実世界2タスクで自律性能が30パーセントポイント以上向上。 - BlenDAgger はポリシー制御と人間介入の間の遷移が57%滑らかになり、訓練データとの軌道類似性が14%高くなる。 - ユーザスタディ (n=14) では、データ収集がより速く (BF=13.32)、主観的知覚に差は見られなかった。

3. 技術・手法の肝は?

- 共有制御を用いて、介入中にポリシーの行動とデモンストレータの行動をブレンドする。 - これにより、完全なテレオペレーションによる介入からポリシーを微調整する従来法と比べて、自律性能が向上する。 - 具体的なブレンドの方法やアルゴリズムの詳細は要旨からは不明。

4. どうやって有効だと検証した?

- 5つの操作タスク (実世界2、シミュレーション3) で評価。 - 実世界2タスクで HG-DAgger と比較して自律性能が30パーセントポイント以上向上。 - 遷移の滑らかさが57%向上、軌道類似性が14%向上。 - ユーザスタディ (n=14) でデータ収集が速いことを確認 (BF=13.32)。

5. 議論はある?

- BlenDAgger は完全にテレオペレーションされた介入からポリシーを微調整する典型的な方法と比較して、より高い自律性能をもたらす。 - ユーザスタディでは主観的知覚に差は見られなかったが、データ収集の速さが利点。 - 議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- HG-DAgger (Human-Gated DAgger) が比較対象として挙げられている。 - その他の関連手法や参照論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

分類: cs.RO, cs.LG

原文アブストラクト

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.

関連論文

PR本紙発行元 EmplifAI