日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.30462

方策校正DAgger:模倣学習のためのオフライン校正ノイズ注入

Policy-Calibrated DAgger: Offline Calibrated Noise Injection for Imitation Learning

シェア:XThreadsFacebookLINEはてブBluesky

生成的な拡散方策の予測行動分布を利用して、専門家軌跡上での方策の不確実性をオフラインで推定し、適切なノイズを注入してDAggerを改善する手法を提案。

詳しい要約

1. どんなもの?

- 本論文は、模倣学習におけるcovariate shift問題に対処するための新しい手法「Policy-Calibrated DAgger」を提案している。 - この手法は、生成的なポリシー(特にdiffusion policy)の特性を利用し、オフラインでポリシーのノイズを推定する。 - 具体的には、専門家の軌跡上の観測において予測される行動の分布の広がりを測定し、記録された軌跡に対する閉ループ誤差を測定する。 - マルチモーダルな行動空間での誤差測定の問題に対処するため、部分的なdenoisingを用いて閉ループ制御中にポリシーを軌跡に誘導し、diffusion modelの特性を利用して誘導しなかった場合の誤差にunnormalizeする。 - ロボットが clutter で狭い環境でエンジンレバーに到達するタスクを設定し、3Dフォトリアリスティックシミュレータと2Dプラナーリーチャー環境で実験を行っている。

2. 先行研究と比べてどこがすごい?

- 既存の手法は、ポリシーが失敗する、または失敗しそうな場所で追加データを収集することでcovariate shiftを緩和するが、一つはロボットを危険な状態に置き、もう一つは適切なノイズ分布を選択する必要がある。 - 提案手法は、最近の生成ポリシーの特性を利用して、ポリシー自身の予測行動分布からオフラインでノイズを推定するため、追加のデータ収集やノイズ分布の選択を必要としない。 - 実験結果では、ノイズなしのdataset aggregationで訓練されたポリシーを上回り、後知恵で最良のノイズレベルと同等の性能を達成し、ノイズレベルのスイープを必要としない。

3. 技術・手法の肝は?

- 提案手法の肝は、diffusion policyの予測行動分布の広がりを利用して、オフラインでポリシーのノイズを推定することである。 - 専門家軌跡上の観測で予測行動の分散を測定し、閉ループ誤差を測定する。 - マルチモーダル行動空間での誤差測定の問題に対処するため、部分的なdenoisingを用いてポリシーを軌跡に誘導し、diffusion modelの特性を利用して測定誤差をunnormalizeする。 - これにより、誘導なしの誤差を推定し、適切なノイズレベルを決定する。

4. どうやって有効だと検証した?

- ロボットが clutter で狭い環境でエンジンレバーに到達するタスクを設定し、3Dフォトリアリスティックシミュレータと2Dプラナーリーチャー環境で実験を行った。 - 提案手法が、ノイズなしのdataset aggregationで訓練されたポリシーを上回り、後知恵で最良のノイズレベルと同等の性能を達成することを示した。 - ノイズレベルのスイープを必要としないことも示した。

5. 議論はある?

- 要旨からは、提案手法の限界や議論についての明示的な記述はない。 - ただし、マルチモーダル行動空間での誤差測定の問題に対処するために部分denoisingとunnormalizeを用いている点が技術的な課題として挙げられる。 - また、実験は特定のタスクと環境に限定されているため、一般性についてはさらなる検証が必要かもしれない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究として、dataset aggregation (DAgger) が挙げられる。 - また、diffusion policy や generative policies が関連手法として言及されている。 - 同分野の定番として、imitation learning における covariate shift 対策の手法(例: DAgger, DART, など)が次に読むべき論文として考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jenny Wang, George Kantor

分類: cs.RO

原文アブストラクト

Policies trained with imitation learning can accumulate errors over time, causing the robot to drift outside the training distribution. Existing methods mitigate this covariate shift by collecting additional data where the policy fails or is likely to fail. The first places the robot in unsafe conditions and the second requires choosing an appropriate noise distribution to collect new expert demonstrations under that noise. We propose Policy-Calibrated DAgger, a method that makes use of the properties of recent generative policies to estimate the policy's noise offline by using its own predicted action distribution. We measure a diffusion policy's spread of predicted actions at observations along the expert trajectory and measure its closed-loop error relative to a recorded trajectory. To address issues with measuring error in a multimodal action space, we guide the policy towards the trajectory during closed-loop control through partial denoising, and use properties of a diffusion model to unnormalize the measured error as if we did not guide it. We experiment in a scenario where a robot is tasked to reach an engine lever in a cluttered and narrow environment and show results in a 3D photorealistic simulator and a 2D planar reacher environment. We show that our method surpasses policies trained with dataset aggregation without noising and matches the performance of the best noise level in hindsight, without requiring a sweep over noise levels.

関連論文

PR本紙発行元 EmplifAI