日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転/VLA/協調運転arXiv:2608.07621v1

CMU-DriveとV2V-VLA:推論ベンチマークと車車間ビジョン・言語・行動モデルによる協調型マルチエージェント統合運転

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

協調自動運転のための閉ループベンチマークCMU-Driveと、車車間通信を統合したVLAモデルV2V-VLAを提案し、複数車両の協調認識・推論・計画を評価する。

詳しい要約

1. どんなもの?

CMU-Driveは、複数のConnected Autonomous Vehicles (CAVs)が安全クリティカルなシナリオで協調運転するためのクローズドループのエンドツーエンドベンチマークであり、V2V-VLAは、協調運転を単一のフォワードパスに統合する協調VLAモデルである。V2V-VLAは、運転アクション、将来のウェイポイント、言語推論、通信ポリシーを同時に生成する。

2. 先行研究と比べてどこがすごい?

既存のVLAモデルは単一の自動運転エージェント向けに設計されており、協調知覚、推論、プランニングのサポートが限定的である。CMU-Driveは、複数のCAVが関与する協調運転を評価する初のクローズドループベンチマークを提供し、V2V-VLAは協調運転を単一のモデルで統合する点で新しい。

3. 技術・手法の肝は?

V2V-VLAは、協調運転を単一のフォワードパスに統合し、運転アクション、将来のウェイポイント、言語推論、通信ポリシーを同時に生成する。CMU-Driveは、安全クリティカルなシナリオと背景交通参加者を含むクローズドループのエンドツーエンドベンチマークを提供する。

4. どうやって有効だと検証した?

CMU-Drive上での実験により、協調VLA運転の最初のベンチマークとベースラインを確立し、その有効性を検証した。

5. 議論はある?

要旨からは、議論や限界についての具体的な記述は不明。ただし、ベンチマークとモデルの公開により、今後の研究の基盤を提供するとしている。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、VLAモデルや協調運転に関する既存研究が関連する。具体的には、Vision-Language-Actionモデルやマルチエージェント協調運転に関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hsu-kuang Chiu, Stephen F. Smith

分類: cs.AI, cs.CV, cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.