日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAナビゲーションarXiv:2610.07558

見えない危険を可視化:温度・放射線対応VLAナビゲーションのための物理誘導視覚プロンプティング

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

シェア:XThreadsFacebookLINEはてブBluesky

凍結したVLAモデルに仮想障害物を重ねるだけで、RGBカメラでは見えない放射線や温度の危険を回避させるプラグアンドプレイ手法を提案し、実機でも検証した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルを用いた Vision-and-Language Navigation (VLN) において、RGB カメラでは検知できない放射線や温度上昇などの不可視リスクを回避するための **Physics-Guided Visual Prompting (PG-VP)** を提案。 - 凍結した VLA モデルを再利用し、危険源の方向に応じて仮想障害物を連続フレームに重畳する **Dynamic Visual Prompting** を行う plug-and-play なマルチモーダル知覚モジュール。 - 危険の種類によらず同一の仮想障害物パターンを用いるため、センサ追加時も視覚プロンプトは固定。危険がなければ何も描画せず、ポリシーは PG-VP なしと同一に振る舞う。

2. 先行研究と比べてどこがすごい?

- 従来の VLA ベース VLN は RGB カメラで見える障害物回避に強みを持つが、放射線・温度などの不可視リスクには対応できない。 - 各リスクに対処するには新しいエンコーダ、データ、再学習が必要で高コストだった。 - PG-VP は凍結 VLA を再学習せずに、既存の「見える障害物回避」能力を再利用して不可視リスクを回避させる点が新しい。 - 危険タイプが変わっても同一の仮想障害物を使うため、センサ追加に対して視覚プロンプトパターンが固定で済む。

3. 技術・手法の肝は?

- 近接する放射線源または熱源が与えられると、物理に基づくリスク評価 (physics-guided risk assessment) を行い、回避方向を決定。 - その方向に対応する仮想障害物を連続フレーム上に重畳する **Dynamic Visual Prompting** を実行。 - ナビゲーションポリシーはこの仮想障害物を通常の障害物と同様に扱い、自然に迂回する。 - 危険の種類に依存せず同一の仮想障害物パターンを使用。危険検出時のみ描画し、非検出時は何も描画しない。

4. どうやって有効だと検証した?

- OmniNav 上で R2R-CE と RxR-CE の val-unseen スプリットを用いて評価。 - 意図した低リスク行動へ導いた割合は 84.9% (R2R-CE) と 83.2% (RxR-CE)。 - ナビゲーション成功率のコストはそれぞれ 6.8 ポイントと 7.9 ポイント。 - 実ロボットで実際の熱源と放射線源を用いた異なるシナリオを再学習なしでテスト。 - 最悪 10% の平均軌道安全性が熱源に対して 63.45%、放射線源に対して 32.59% 改善。

5. 議論はある?

- ナビゲーション成功率の低下 (6.8–7.9 ポイント) と安全性向上のトレードオフが示唆される。 - 危険タイプによらず同一の仮想障害物を使う設計の一般性・限界については要旨からは不明。 - 実ロボット実験での環境条件や危険源の強度など詳細は要旨からは不明。 - 他の不可視リスク (例: 化学物質、音響) への拡張性は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている **OmniNav**、**R2R-CE**、**RxR-CE** に関する研究。 - 関連手法として **Vision-Language-Action (VLA)** モデル、**Vision-and-Language Navigation (VLN)**、**Dynamic Visual Prompting** の元となった visual prompting 研究。 - 比較対象として、不可視リスクごとに専用エンコーダ・データ・再学習を要する従来の安全クリティカル施設向けナビゲーション手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hojoon Son, Fan Zhang

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and model retraining. We propose Physics-Guided Visual Prompting (PG-VP), a plug-and-play multimodal perception module that instead reuses what a frozen VLA model already does well: avoiding visible obstacles. Given a proximal radiation or thermal source, PG-VP performs a physics-guided risk assessment to determine the avoidance direction and overlays a corresponding virtual obstacle that moves across consecutive frames (Dynamic Visual Prompting). The navigation policy then naturally detours around this invisible hazard. The identical virtual obstacle is used regardless of hazard type, so the visual prompting pattern remains fixed as sensors are added. When no hazard is detected, nothing is rendered, and the policy behaves exactly as it would without PG-VP. We evaluate PG-VP on OmniNav using the val-unseen splits of R2R-CE and RxR-CE, where it guides the policy toward intended low-risk actions in 84.9% and 83.2% of cases, at a cost of 6.8 and 7.9 percentage points in navigation success rate. We further test it with distinct scenarios on a real robot in the presence of actual thermal and radiation sources, all without any retraining. The real test shows that PG-VP effectively avoids these invisible hazards, improving worst-10% average trajectory safety by 63.45% and 32.59% against thermal and radiation sources, respectively.

関連論文

PR本紙発行元 EmplifAI