言語条件付き走行可能性表現の学習による適応的視覚ナビゲーション
Learning Language-Conditioned Traversability Representations for Adaptive Visual Navigation
視覚言語モデルで言語指示に応じた走行可能性とゴールを推定し、高速なフローマッチングプランナで非同期に経路計画するナビゲーション手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Senda Chen, Changxu Cheng, Fangdi Li, Tao Wang, Wuyue Zhao
分類: cs.RO
原文アブストラクト
Traversability is essential for visual navigation but varies with robot capabilities and user preferences. Conventional pipelines often rely on explicit costmaps or segmentation masks with predefined criteria, requiring hand-crafted rules and careful tuning. Moreover, viewpoint-dependent segmentation masks complicate asynchronous planning under perception latency. We present LaTraNav, a framework that learns language-conditioned traversability representations for adaptive visual navigation. Its asynchronous architecture combines a slow vision-language model that produces latent representations of traversability and navigation goals, with a fast flow-matching planner conditioned on these representations. To train the system, we develop a simulation-based data generation pipeline with controllable trajectories, producing observations paired with language instructions, traversability maps, goal locations, and diverse trajectories. Photorealistic image translation further enhances visual realism. Evaluations on datasets from multiple sources demonstrate effective language-guided traversability segmentation and goal localization by the slow VLM, alongside adaptive pixel-space path planning by the fast planner. Latent conditioning improves planning performance over explicit segmentation masks, while asynchronous scheduling increases the path-update rate by $6.05\times$ at the same semantic-update rate.
関連論文
- WAND: 複雑な風外乱と密集障害物下での四 rotor のロバストナビゲーション学習ナビゲーション
- 信念に基づく行動と必要時の知覚:断続的知覚下のナビゲーションのためのベイズ空間世界モデルナビゲーション
- 経路創出型ナビゲーション:身体性インタラクションによるロボットナビゲーションナビゲーション
- LiteNWM: 実環境でのオンボード視覚ナビゲーションのための効率的な潜在世界モデルナビゲーション
- SuperNav: あらゆるシーンであらゆるタスクに対応するエージェント型ナビゲーションシステムナビゲーション
- STAG: グリッドベースコストマップからの疎な走行性考慮グラフ表現によるロボットナビゲーションナビゲーション