日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2610.09695

優れた視覚表現は常にエンドツーエンド自動運転を改善するのか?

Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

シェア:XThreadsFacebookLINEはてブBluesky

視覚基盤モデルの表現をプランナーに整合させるViRAを提案し、VFM表現が運転性能を一貫して改善することを示した。

詳しい要約

1. どんなもの?

- 視覚基盤モデル(VFM)を end-to-end 自動運転に統合する際、その表現が運転性能をいつ向上させるかを調べる研究。 - planner 非依存の visual representation alignment フレームワーク ViRA を導入。 - planner のアーキテクチャと推論コストを変えずに表現を整列。 - 知見に基づき ViRA-Diffusion を開発。

2. 先行研究と比べてどこがすごい?

- VFM 表現が常に運転性能を改善するとは限らない点を検証。 - 多様な end-to-end planner で一貫した改善を示し、zero-shot closed-loop 評価にも利得が及ぶ。 - VFM target の選択が planning 性能に影響し、異なる VFM への整列が事前学習済み VFM encoder を持つ planner にさらに利する場合がある。 - 補助知覚 supervision が VFM target 選択への感度を下げ、EPDMS のばらつきを 2.7 から 0.5 点に縮小。

3. 技術・手法の肝は?

- planner アーキテクチャと推論コストを固定した planner-agnostic な visual representation alignment。 - VFM を target として視覚表現を整列。 - 補助知覚 supervision の有無を比較。 - 知見を踏まえ、補助知覚 supervision なしで訓練する diffusion-based planner ViRA-Diffusion を開発。

4. どうやって有効だと検証した?

- 多様な end-to-end planner で VFM-guided 視覚表現の改善を確認。 - zero-shot closed-loop 評価へ利得が及ぶことを確認。 - 5 つの target 間の EPDMS ばらつきを補助知覚 supervision ありで 2.7 から 0.5 点に縮小。 - ViRA-Diffusion が NAVSIM v2 navtest で 92.3 EPDMS を達成し、比較した最近手法を少なくとも 1.9 点上回る。

5. 議論はある?

- VFM 統合時には target 選択と planner supervision を併せて検討する必要がある。 - 補助知覚 supervision は効果の低い VFM target を補償しうる。 - 要旨からは不明:具体的な失敗条件や限界、計算コストの詳細。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:ViRA、ViRA-Diffusion、NAVSIM v2。 - 関連手法:end-to-end autonomous driving、visual foundation models (VFMs)、diffusion-based planner、EPDMS。 - 同分野の定番:closed-loop evaluation、planner-agnostic alignment。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zihao Zhang, Haochen Tian, Tianyu Li, Changhui Jing, Jingliang He, Naisheng Ye, Ziyuan Pu, Zhenjie Yang

分類: cs.RO, cs.CV

原文アブストラクト

Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at https://github.com/OpenDriveLab/ViRA.

関連論文

PR本紙発行元 EmplifAI