VIGOR: モデルベース強化学習における潜在空間一貫性によるゼロショット視覚汎化
VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
モデルベース強化学習において、弱から強への非対称なデータ拡張と潜在空間での動力学・エンコーダ一貫性を組み合わせ、未知の視覚的妨害に対してゼロショットで汎化できるフレームワークを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee
分類: cs.AI, cs.LG, cs.RO
原文アブストラクト
Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.
関連論文
- 行動なし時系列からの動的埋め込みによる転移可能な方策学習モデルベース強化学習
- CEMにおける世界モデルは提案メカニズムでもあるモデルベース強化学習
- 計画と学習のループを閉じる:学習済み世界モデルによるロボット制御モデルベース強化学習
- 速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習モデルベース強化学習
- 表現世界モデル:表現空間における状態・遷移・実行可能計画の学習モデルベース強化学習
- CAST: 交互状態価値目標と拡張方策勾配によるモデルベース強化学習モデルベース強化学習