Geo-VLA: 地図セマンティクスの内在化による幾何認識型視覚言語行動計画
Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
複雑な運転環境でのVLAモデルの計画性能を向上させるため、幾何学的な地図情報を学習時に内在化するプラグアンドプレイ型フレームワークGeo-VLAを提案した。推論時にはHD地図を不要とし、NAVSIM v1で単眼カメラVLAプランナーとして最高性能を達成した。
著者: Ran Chen, Jiaxing Ren, Zhikun Zhang, Yunhao Hou, Junbao Zhuo, Bochao Zou
分類: cs.RO, cs.AI
原文アブストラクト
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations. During training, Geo-VLA internalizes geometric map semantics to strengthen road-structure representations, while requiring no HD maps or additional lane information during inference. To support this approach, we introduce Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments on NAVSIM v1 demonstrate that Geo-VLA consistently improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and establishing a new state-of-the-art among single-camera VLA planners.