日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モデルベース強化学習arXiv:2610.02801

VIGOR: モデルベース強化学習における潜在空間一貫性によるゼロショット視覚汎化

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

モデルベース強化学習において、弱から強への非対称なデータ拡張と潜在空間での動力学・エンコーダ一貫性を組み合わせ、未知の視覚的妨害に対してゼロショットで汎化できるフレームワークを提案した。

詳しい要約

1. どんなもの?

- モデルベース強化学習(MBRL)における視覚的汎化を目指すフレームワーク VIGOR を提案。 - 学習済み潜在ダイナミクス内で計画する MBRL は、未見の視覚的妨害(背景変化、照明変化、カメラシフト)下で性能が大幅に低下する問題に対処。 - ゼロショットで未見の視覚的妨害に汎化しつつ、MBRL のサンプル効率を維持する。 - 3つの相互依存コンポーネント:非対称 weak-to-strong augmentation、dynamics-level consistency、encoder-level stabilization を統合。

2. 先行研究と比べてどこがすごい?

- モデルフリー RL ではエンコーダ摂動は単一ステップ予測のみに影響するが、MBRL では視覚的妨害がエンコーダ出力を分布外に押し出し、その誤差が計画ホライズンにわたる再帰的潜在ロールアウトで複合する二段階の脆弱性がある。 - VIGOR はこの二段階脆弱性に対処し、ゼロショット汎化を実現。 - DeepMind Control Suite (DMC) と Robosuite で、最先端のモデルフリーおよびモデルベースのベースラインを上回り、2番目に良いベースラインを DMC で 3.4%、Robosuite で 43.6% 上回る。

3. 技術・手法の肝は?

- 非対称 weak-to-strong augmentation:単一バッチ内で weak-only と weak-to-strong の潜在ビューをペアにする。 - dynamics-level consistency:直接潜在回帰を通じて、拡張不変の遷移予測を強制する。 - encoder-level stabilization:dynamics-level consistency による cross-augmentation 監督下でエンコーダのドリフトを防ぐ。 - これら3要素が相互依存し、潜在空間の一貫性を実現。

4. どうやって有効だと検証した?

- DeepMind Control Suite (DMC) と Robosuite で評価。 - 最先端のモデルフリーおよびモデルベースのベースラインと比較し、VIGOR が上回ることを示す。 - アブレーションにより、VIGOR のロバスト性が拡張に依存しない(augmentation-agnostic)ことを確認:デフォルト拡張を異なる摂動ファミリーの代替に置き換えても強い汎化が保たれ、ロバスト性を駆動するのは拡張の選択ではなく潜在空間の一貫性であることを確認。

5. 議論はある?

- アブレーションから、潜在空間の一貫性がロバスト性の主要因であり、拡張の選択は本質的でないことが示唆される。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:モデルフリー RL、モデルベース RL(MBRL)のベースライン。 - 関連手法:DeepMind Control Suite (DMC)、Robosuite を用いた研究。 - 具体的な論文名は要旨からは不明。同分野の定番として、Dreamer、PlaNet、CURL、DrQ などが考えられるが、要旨に明示なし。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee

分類: cs.AI, cs.LG, cs.RO

原文アブストラクト

Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.

関連論文

PR本紙発行元 EmplifAI