日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2610.04054

Reward-DAgger: 汎用進捗報酬モデルによるロボット主導の対話型模倣学習

Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

シェア:XThreadsFacebookLINEはてブBluesky

汎用報酬モデルが出力する進捗信号を利用して、人間の介入が必要なタイミングをロボット自身が判断する対話型模倣学習フレームワークを提案し、タスクごとの再調整なしに失敗検出と回復行動の学習を可能にした。

詳しい要約

1. どんなもの?

- ロボット学習における一般的な制御ポリシーの性能低下を検出し、回復行動を教えるためのフレームワーク。 - Reward-DAggerは、汎用報酬モデルからの密な進捗信号を利用して、人間の介入が必要なタイミングを決定するロボットゲート型の対話型模倣学習。 - ポリシーアーキテクチャに依存せず、ポリシー内部へのアクセスを必要とせず、タスク間でゲーティングメカニズムを再調整せずに適用可能。

2. 先行研究と比べてどこがすごい?

- 既存のランタイムモニタリング手法は、タスクやポリシーに特化した訓練やハイパーパラメータ調整を必要とし、クロスタスク展開を制限し、反復的なポリシー更新時に追加のオーバーヘッドを生じる。 - Reward-DAggerは、タスクやポリシーに依存しない汎用報酬モデルを利用し、再調整なしでクロスタスク展開を可能にする。 - 失敗検出の精度と遅延のトレードオフにおいて、既存のランタイムモニタリングベースラインよりも優れる。

3. 技術・手法の肝は?

- 汎用報酬モデルから得られる密な進捗信号を利用して、人間の介入が必要なタイミングを決定するロボットゲート型対話型模倣学習フレームワーク。 - ポリシーアーキテクチャに非依存で、ポリシー内部にアクセスせず、タスク間でゲーティングメカニズムを再調整する必要がない。 - 具体的な報酬モデルの詳細やゲーティングの実装は要旨からは不明。

4. どうやって有効だと検証した?

- 8つのシミュレーションおよび実世界タスクで評価。 - Reward-DAggerは対話型学習を通じて下流ポリシーの自律成功率を一貫して改善し、人間の労力に対する強いリターンを達成し、ほとんどの設定でベースラインを上回った。 - 同じゲーティング設定をタスク間で使用し、タスク固有のハイパーパラメータ調整なしで、タスク、環境、ポリシーアーキテクチャ間の転移を示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。同分野の定番として、DAgger (Dataset Aggregation) やInteractive Imitation Learning、Runtime Monitoring、General-Purpose Reward Models などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ryan Li, Yigit Korkmaz, Erdem Bıyık

分類: cs.RO, cs.AI

原文アブストラクト

Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.

関連論文

PR本紙発行元 EmplifAI