日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデル/ロバスト制御arXiv:2610.07599

世界モデルにおける潜在外乱のモデリングによるロバスト意思決定

Modeling Latent Disturbances for Robust Decision-Making in World Models

シェア:XThreadsFacebookLINEはてブBluesky

世界モデルの潜在空間に、もっともらしいが悲観的な遷移を引き起こす外乱を定義し、conformal predictionで不確実性集合を調整してロバストな行動選択を実現する手法を提案。

詳しい要約

1. どんなもの?

- World Models (WMs) の latent space 上で robust decision-making を行う研究。 - 明示的な dynamics と物理的 disturbance が与えられる通常の robust optimization を、高次元観測から学習した WM の latent space に適用する際の課題を扱う。 - latent-space disturbance を学習済み latent dynamics への摂動としてモデル化し、悲観的だが plausible な遷移を誘導する。 - ロボット行動と最悪ケース latent disturbance を game-theoretic optimization で同時最適化する。

2. 先行研究と比べてどこがすごい?

- 従来の robust optimization は dynamics が明示的に指定され、物理的に意味のある disturbance を前提とする。 - WM は state space と dynamics を高次元観測から学習するため、latent-space disturbance の定義が不明確という根本的課題があった。 - 本研究は dynamics-aware similarity metric と out-of-distribution detection を組み合わせ、plausible な latent dynamics 集合を構築する点が新しい。 - conformal prediction で uncertainty set を calibration し、過度に悲観的にならないようにする点が先行研究と異なる。

3. 技術・手法の肝は?

- latent-space disturbance を学習済み latent dynamics への摂動として定式化する。 - dynamics-aware similarity metric で plausible な遷移を捉え、out-of-distribution detection で implausible な latent state を除外する。 - conformal prediction を用いて latent dynamics 上の uncertainty set を calibration する。 - game-theoretic optimization により robust なロボット行動と最悪ケース latent disturbance を同時最適化する。 - latent safety filtering と generative control policy の sample-and-verify steering の2つのパラダイムに適用する。

4. どうやって有効だと検証した?

- 制御された simulation 実験で、latent disturbance が WM latent space 上で robust decision-making を可能にすることを示す。 - Franka manipulator を用いた hardware 実験を実施。 - safety filtering では失敗を70%削減、sampling-based policy steering では54%削減した。 - プロジェクトサイト: https://junwon.me/LatentDisturbance/

5. 議論はある?

- latent disturbance が悲観的すぎず plausible に保たれるよう conformal prediction で calibration する点が議論の焦点。 - 2つのパラダイム(latent safety filtering と sample-and-verify steering)への適用可能性を検討。 - 詳細な限界や議論の内容は要旨からは不明。

6. 次に読むべき論文は?

- World Models (WMs) に関する研究 - Robust optimization - Conformal prediction - Out-of-distribution detection - Game-theoretic optimization - Generative control policy - 要旨で参照/比較されている個別の研究は明示されていないため、上記の関連手法・分野を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junwon Seo, Andrea Bajcsy

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: https://junwon.me/LatentDisturbance/.

PR本紙発行元 EmplifAI