ProxPI: 学習済み事前分布の不一致下でのサンプリングベースMPCに対する近接事前分布注入
ProxPI: Proximal Prior Injection for Sampling-Based MPC under Learned-Prior Mismatch
学習済みポリシーとMPCを組み合わせる際、事前分布が不正確な場合にサンプリング分布をポリシー出力に中心化する既存手法の問題を指摘し、名目中心のサンプリングを維持しつつポリシーを近接コストとして組み込むProxPIを提案。理論と実験で性能回復を実証。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Euncheol Im, Myotaeg Lim, Yisoo Lee
分類: cs.RO, eess.SY
原文アブストラクト
Combining learned policies with model predictive control can leverage learned task priors while retaining online adaptation to new objectives and constraints, but performance degrades when the policy is out of distribution. In policy-guided model predictive path integral (MPPI) control, a policy-centered warm-start approach centers the sampling distribution on the policy output. When the prior is mismatched, centering the sampling distribution on the policy output restricts exploration around an unsuitable solution and prevents recovery toward the task optimum. We propose Proximal Prior Injection (ProxPI), which retains nominal-centered MPPI sampling and incorporates the policy through a soft proximity cost. This matches the in-distribution performance of existing prior-injection schemes while enabling the optimizer to escape an inaccurate policy and recover vanilla MPPI-level performance. We theoretically show that re-centering on the prior discards the optimizer's correction at every update, whereas nominal-centered sampling retains it and converges to a solution set by both the task cost and the prior, and that this failure is not removed by a larger rollout budget. Simulations and real-robot experiments demonstrate robust performance under both in-distribution and out-of-distribution tasks.