日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
安全フィルタarXiv:2609.34300

世界モデルが嘘をつくとき:誤った想像下での適応的安全解析

When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations

シェア:XThreadsFacebookLINEはてブBluesky

世界モデルの予測誤差を直接観測し、適応的共形推論で不確実性集合を構築することで、潜在空間安全フィルタを較正する手法を提案。

詳しい要約

1. どんなもの?

- 高次元ロボットシステムの安全性推論にWorld Modelを活用する研究。 - World Modelの予測誤差(バイアス、不確かさ、自信過剰な誤り)に対処する適応的潜在安全フィルタを提案。 - 観測から直接推定したWorld Model誤差を用いて安全推論を校正する。

2. 先行研究と比べてどこがすごい?

- 既存の潜在安全フィルタはensemble disagreementやvalue-target consistency residualsなどの補助信号に依存。 - これらの信号はWorld Modelの予測が観測から逸脱しても小さいままである可能性がある。 - 提案手法は直接観測されたWorld Model誤差を利用し、より信頼性の高い適応を実現。

3. 技術・手法の肝は?

- Adaptive Conformal Inferenceを用いて、予測潜在状態と観測推論潜在状態の不一致からオンライン不確かさ集合を構築。 - 学習済み価値関数をこの集合上で最小化することで悲観的に安全性を評価。 - World Modelが正確なときは保守性を最小限に抑え、不一致が観測されると cautious になる。 - 適応的不確かさ半径の有限時間カバレッジ保証を提供。

4. どうやって有効だと検証した?

- シミュレーションとハードウェア実験を実施。 - 最先端の潜在安全フィルタと比較して故障を大幅に削減し、タスク完了を維持することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:ensemble disagreement、value-target consistency residuals、latent safety filters、Hamilton-Jacobi safety value functions、Adaptive Conformal Inference。 - 関連手法:World Models、Safe Reinforcement Learning、Conformal Prediction。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: John Cao, Somil Bansal

分類: cs.RO, cs.AI, eess.SY

原文アブストラクト

World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.

関連論文

PR本紙発行元 EmplifAI