日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
V2X/協調運転/VLMarXiv:2608.21032v1

路側協調自動運転:データ基盤から視覚言語エンドツーエンド推論へ

Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

V2X協調運転のためのシミュレーションプラットフォームとVQAデータセットを構築し、視覚言語モデルを用いたエンドツーエンドの協調運転フレームワークAURORAを提案した。閉ループ評価で高い性能を達成した。

著者: Yitao Xu, Tong Wu, Yiyan Wu, Guoji Xu, Yanbo Jiang, Jiahao Wang, Zehong Ke, Junkai Jiang, Fang Zhang, Jianqiang Wang

分類: cs.RO

原文アブストラクト

Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-end cooperative driving. To address these limitations, we introduce V2XBench, a simulation platform featuring synchronized ego--roadside sensing and closed-loop evaluation, together with Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning. Building upon this benchmark infrastructure, we propose AURORA, an end-to-end cooperative driving framework. Equipped with a dual-view perception architecture, AURORA mitigates spatial and semantic discrepancies across ego and roadside viewpoints through a query-level Cross-View Query Alignment and Fusion (CQAF) module. Leveraging the resulting unified tokens, a LoRA-adapted VLM bridges semantic reasoning and generative trajectory planning. Extensive closed-loop evaluations on V2XBench demonstrate that AURORA achieves state-of-the-art performance in heavily occluded scenarios, with a Route Completion rate of 98.21% and a Driving Score of 76.02, while requiring low roadside communication bandwidth. Ultimately, this work pioneers an extensible V2X--VLM paradigm, paving the way for next-generation cooperative autonomous driving.