日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2610.02765

自動運転のための視覚言語モデルを用いた局所適合安全モニタリング

Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

凍結した視覚言語モデルの予測を、シーンに応じた局所的な適合予測で確率的に校正し、衝突を引き起こす軌道の検出率を大幅に向上させた。

詳しい要約

1. どんなもの?

- 自動運転の計画軌道に対する衝突リスクを、Vision-Language Models (VLMs) で監視する研究。 - 提案手法 Split Label-Localized Conformal Prediction (SLLCP) は、凍結した VLM の予測を確率的に校正された安全予測集合へ変換する post-hoc 校正層。 - 観測された走行シーンに応じて安全推定能力が変わる点を考慮し、不確実性閾値の計算で関連する過去経験を重み付けする localized 手続きを導入。 - exchangeability の下で label-conditional な有限サンプル分布フリー被覆を提供。

2. 先行研究と比べてどこがすごい?

- 既存の古典的アプローチは予測モデルの品質に制限されることが多い。 - VLM は高レベル行動の結果推論に有望だが、その近似予測は自動運転のような安全クリティカル用途には不適切。 - Conformal prediction (CP) は black-box モデル予測の不確実性を定量化するデータ駆動フレームワークとして登場。 - SLLCP は凍結 VLM 上に post-hoc 校正層を設け、信頼できない予測を確率的に校正された安全予測集合に変換する点が新しい。

3. 技術・手法の肝は?

- Split Label-Localized Conformal Prediction (SLLCP) を提案。 - 凍結した VLM の上に post-hoc 校正層を適用。 - 観測された走行シーンに依存して安全推定能力が変わることを考慮し、localized 手続きで関連する過去経験を upweight して不確実性閾値を計算。 - exchangeability の下で label-conditional な有限サンプル分布フリー被覆を提供。

4. どうやって有効だと検証した?

- 未見シナリオの CARLA 軌道 15k で評価。 - SLLCP は衝突を引き起こす軌道を Qwen backbone で 89.6%、Cosmos backbone で 88.4% 正しくフラグ。 - ベース VLM はそれぞれ 4.6% と 39.1% しかフラグできなかった。 - これらの結果は、local かつ label-conditional な校正が危険な軌道の見逃しを減らせることを示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: Conformal prediction (CP)、Vision-Language Models (VLMs)、Qwen、Cosmos、CARLA。 - 関連手法として、Split Conformal Prediction、Label-Conditional Conformal Prediction、Localized Conformal Prediction などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Luís Marques, Rong Fang, Disha Kamale, Dmitry Berenson

分類: cs.RO, cs.LG

原文アブストラクト

Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the ego motion. Existing classical approaches are often limited by the quality of their forecasting model. Vision-language models (VLMs) have shown promise in reasoning about the consequences of high-level actions, yet their approximate predictions are unsuitable for safety-critical applications such as autonomous driving. Conformal prediction (CP) has emerged as a data-driven framework for quantifying the uncertainty of black-box model predictions. We propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen VLMs that transforms their unreliable predictions into probabilistically calibrated safety prediction sets. We consider how the ability to estimate safety can depend on the observed driving scene and introduce a localized procedure that upweights relevant past experience when calculating uncertainty thresholds. We provide label-conditional finite-sample distribution-free coverage under exchangeability. Evaluated over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, while the base VLMs only flagged 4.6% and 39.1% of the collision-causing trajectories, respectively. These results indicate that local, label-conditional calibration can reduce missed unsafe trajectories.

関連論文

PR本紙発行元 EmplifAI