Talk2Escape: 対話による視覚言語ナビゲーションの閉ループ化
Talk2Escape: Conversational Grounding for Vision-and-Language Navigation
ナビゲーションエージェントの迷走を検知したら対話で修正指示を求める仕組みを提案し、シミュレータと実機で成功率と頑健性を向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Zerui Li, Sihao Lin, Yanyan Shao, Jiwen Zhang, Xiangyu Shi, Shijie Li, Qi Wu
分類: cs.RO, cs.HC
原文アブストラクト
While Vision-and-Language Navigation (VLN) has demonstrated remarkable success, the prevailing single-turn paradigm exposes a fundamental vulnerability: agents operate in a strictly open-loop manner. In practice, factors such as perceptual aliasing, sensor noise, and odometry drift can cause minor deviations to accumulate over time, often leading to catastrophic mission failures with no built-in mechanism for error recovery. To address this, we introduce \textit{Talk2Escape}, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process. At its core, a lightweight vision-language module continuously monitors agent kinematics. Upon detecting localized looping or severe trajectory divergence, it translates raw egocentric observations into concise, grounded queries to solicit targeted corrective feedback from either an algorithmic oracle or a human-in-the-loop. Extensive evaluations in high-fidelity simulators, including R2R-CE, RxR-CE, and VLNVerse, demonstrate that \textit{Talk2Escape} exhibits consistent improvements across diverse base agents. Empirically, \textit{Talk2Escape} achieves a 66.0\% Success Rate on R2R-CE, outperforming the current supervised and zero-shot state-of-the-art methods. We further validate its sim-to-real transfer on a Unitree Go2 quadruped, proving that proactive dialogue drastically improves navigation robustness in physical environments.
関連論文
- FSD-VLN: 空中長距離視覚言語ナビゲーションのための高速・低速デュアルシステムモデリング視覚言語ナビゲーション
- SEDualVLN: 空間強化デュアルシステムによる視覚言語ナビゲーション視覚言語ナビゲーション
- Web動画からの暗黙的幾何表現を用いた視覚言語ナビゲーション視覚言語ナビゲーション
- RAGNav: マルチゴール視覚言語ナビゲーションのための検索拡張トポロジカル推論フレームワーク視覚言語ナビゲーション
- ドリフトを未然に防ぐ:堅牢な視覚言語ナビゲーションのための遡及的修正視覚言語ナビゲーション
- VL-Nav: ニューロシンボリック推論に基づく視覚言語ナビゲーション視覚言語ナビゲーション