日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2610.11622

言語条件付き走行可能性表現の学習による適応的視覚ナビゲーション

Learning Language-Conditioned Traversability Representations for Adaptive Visual Navigation

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルで言語指示に応じた走行可能性とゴールを推定し、高速なフローマッチングプランナで非同期に経路計画するナビゲーション手法を提案。

詳しい要約

1. どんなもの?

- 言語条件付きの走行可能性表現を学習する適応的視覚ナビゲーションのフレームワーク「LaTraNav」を提案。 - 非同期アーキテクチャで、遅いVLMが走行可能性とナビゲーション目標の潜在表現を生成し、速いflow-matchingプランナーがそれに条件付けられて動作。 - シミュレーションベースのデータ生成パイプラインを開発し、言語指示、走行可能性マップ、目標位置、多様な軌跡を含むデータを生成。 - フォトリアリスティックな画像変換で視覚的リアリズムを強化。 - 複数ソースのデータセットで評価し、言語ガイド付き走行可能性セグメンテーションと目標位置特定、適応的ピクセル空間経路計画を実証。

2. 先行研究と比べてどこがすごい?

- 従来のパイプラインは明示的なコストマップやセグメンテーションマスクに依存し、手作りルールと調整が必要。 - 視点依存のセグメンテーションマスクは知覚遅延下での非同期計画を複雑化。 - LaTraNavは言語条件付き潜在表現を学習し、明示的セグメンテーションを不要に。 - 潜在条件付けは明示的セグメンテーションマスクよりも計画性能を向上。 - 非同期スケジューリングにより、同じセマンティック更新レートで経路更新レートを6.05倍に増加。

3. 技術・手法の肝は?

- 非同期アーキテクチャ:遅いVLMが走行可能性とナビゲーション目標の潜在表現を生成し、速いflow-matchingプランナーがそれに条件付けられてピクセル空間経路を計画。 - シミュレーションベースのデータ生成パイプライン:制御可能な軌跡で観測と言語指示、走行可能性マップ、目標位置、多様な軌跡をペアで生成。 - フォトリアリスティック画像変換により視覚的リアリズムを強化。 - 言語条件付き走行可能性表現の学習。

4. どうやって有効だと検証した?

- 複数ソースのデータセットで評価。 - 遅いVLMによる言語ガイド付き走行可能性セグメンテーションと目標位置特定の有効性を実証。 - 速いプランナーによる適応的ピクセル空間経路計画を実証。 - 潜在条件付けが明示的セグメンテーションマスクよりも計画性能を向上させることを確認。 - 非同期スケジューリングが同じセマンティック更新レートで経路更新レートを6.05倍に増加させることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、flow-matchingプランナー、VLM(Vision-Language Model)、言語条件付きナビゲーション、走行可能性セグメンテーション、非同期計画に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Senda Chen, Changxu Cheng, Fangdi Li, Tao Wang, Wuyue Zhao

分類: cs.RO

原文アブストラクト

Traversability is essential for visual navigation but varies with robot capabilities and user preferences. Conventional pipelines often rely on explicit costmaps or segmentation masks with predefined criteria, requiring hand-crafted rules and careful tuning. Moreover, viewpoint-dependent segmentation masks complicate asynchronous planning under perception latency. We present LaTraNav, a framework that learns language-conditioned traversability representations for adaptive visual navigation. Its asynchronous architecture combines a slow vision-language model that produces latent representations of traversability and navigation goals, with a fast flow-matching planner conditioned on these representations. To train the system, we develop a simulation-based data generation pipeline with controllable trajectories, producing observations paired with language instructions, traversability maps, goal locations, and diverse trajectories. Photorealistic image translation further enhances visual realism. Evaluations on datasets from multiple sources demonstrate effective language-guided traversability segmentation and goal localization by the slow VLM, alongside adaptive pixel-space path planning by the fast planner. Latent conditioning improves planning performance over explicit segmentation masks, while asynchronous scheduling increases the path-update rate by $6.05\times$ at the same semantic-update rate.

関連論文

PR本紙発行元 EmplifAI