CAVE-NAV: VLMを用いた水中洞窟環境における自律3次元ナビゲーション
CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments
水中洞窟での自律ナビゲーションのために、視覚言語モデルと思考連鎖推論を用いて環境の手がかりから進行可能な方向を推定し、衝突回避を実現するフレームワークを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Zhenqi Wu, Yuanjie Lu, Yisheng Zhang, Miao Yu, Xuesu Xiao, Jaejeong Shin, Xiaomin Lin
分類: cs.RO
原文アブストラクト
Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.