StageVLN: 空間・軌跡補助ガイダンスによる効率的な視覚言語ナビゲーション
StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation
訓練時のみ幾何学基盤モデルと軌跡監督でナビゲーション表現を強化し、推論時は追加計算なしで高い性能を達成するフレームワーク。
著者: Anh Dao, Quan-Dung Pham, Le Danh Vinh, The Anh Nguyen, Nguyen Viet Tri Pham, Yiyu Chen, Pham Tuyen Le, Van-Truong Nguyen, Quan Nguyen
分類: cs.CV, cs.RO
原文アブストラクト
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3\% SR and 51.4\% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3\% SR without additional navigation training data or a geometry encoder at inference.