日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
視覚言語ナビゲーションarXiv:2609.28296

Talk2Escape: 対話による視覚言語ナビゲーションの閉ループ化

Talk2Escape: Conversational Grounding for Vision-and-Language Navigation

シェア:XThreadsFacebookLINEはてブBluesky

ナビゲーションエージェントの迷走を検知したら対話で修正指示を求める仕組みを提案し、シミュレータと実機で成功率と頑健性を向上させた。

詳しい要約

1. どんなもの?

- Vision-and-Language Navigation (VLN) における単一ターン・オープンループの脆弱性を解決するため、対話による介入フレームワーク Talk2Escape を提案。 - ナビゲーションを閉ループの対話プロセスとして再定義し、エージェントの運動を監視して必要時に修正フィードバックを求める。 - モデル非依存で、アルゴリズム的オラクルまたは人間の介入者からフィードバックを得る。

2. 先行研究と比べてどこがすごい?

- 従来の単一ターン・オープンループ手法は知覚エイリアシングやセンサノイズ、オドメトリドリフトによる誤差蓄積で失敗しやすい。 - Talk2Escape は能動的かつモデル非依存の対話介入により、エラー回復機構を組み込む点が新しい。 - 多様なベースエージェントで一貫した改善を示し、R2R-CE で 66.0% の成功率を達成し、教師あり・ゼロショットの最先端手法を上回る。

3. 技術・手法の肝は?

- 軽量な vision-language モジュールがエージェントの運動を連続監視。 - 局所的なループや深刻な軌道逸脱を検出すると、生の egocentric 観測を簡潔で grounded なクエリに変換。 - アルゴリズム的オラクルまたは人間のループ内介入者から的を絞った修正フィードバックを要請。

4. どうやって有効だと検証した?

- 高忠実度シミュレータ R2R-CE, RxR-CE, VLNVerse で広範に評価。 - 多様なベースエージェントで一貫した改善を確認。 - R2R-CE で 66.0% の成功率を達成し、教師あり・ゼロショットの最先端手法を上回る。 - Unitree Go2 四足歩行ロボットで sim-to-real 転移を検証し、物理環境でのロバスト性向上を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- R2R-CE, RxR-CE, VLNVerse などのベンチマーク。 - 教師あり・ゼロショットの最先端 VLN 手法。 - Unitree Go2 を用いた sim-to-real 研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zerui Li, Sihao Lin, Yanyan Shao, Jiwen Zhang, Xiangyu Shi, Shijie Li, Qi Wu

分類: cs.RO, cs.HC

原文アブストラクト

While Vision-and-Language Navigation (VLN) has demonstrated remarkable success, the prevailing single-turn paradigm exposes a fundamental vulnerability: agents operate in a strictly open-loop manner. In practice, factors such as perceptual aliasing, sensor noise, and odometry drift can cause minor deviations to accumulate over time, often leading to catastrophic mission failures with no built-in mechanism for error recovery. To address this, we introduce \textit{Talk2Escape}, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process. At its core, a lightweight vision-language module continuously monitors agent kinematics. Upon detecting localized looping or severe trajectory divergence, it translates raw egocentric observations into concise, grounded queries to solicit targeted corrective feedback from either an algorithmic oracle or a human-in-the-loop. Extensive evaluations in high-fidelity simulators, including R2R-CE, RxR-CE, and VLNVerse, demonstrate that \textit{Talk2Escape} exhibits consistent improvements across diverse base agents. Empirically, \textit{Talk2Escape} achieves a 66.0\% Success Rate on R2R-CE, outperforming the current supervised and zero-shot state-of-the-art methods. We further validate its sim-to-real transfer on a Unitree Go2 quadruped, proving that proactive dialogue drastically improves navigation robustness in physical environments.

関連論文

PR本紙発行元 EmplifAI