日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.17728

RAF-VLA: 未来表現との整合によるエンドツーエンド自動運転

RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

将来フレームの表現を直接教師信号として方策の内部表現を整列させることで、未来生成を伴わずに自動運転VLAの計画性能を高める手法を提案。

詳しい要約

1. どんなもの?

- 自動運転向けのVision-Language-Action (VLA) モデルであるRAF-VLAを提案。 - 未来の運転シーンを予測するWorld-Modeling VLAの課題(追加の訓練負担と推論遅延)を解決。 - 未来フレーム表現からの直接的なガイダンスにより、計画に関連する内部表現を形成。 - Future-Aligned Supervised Fine-Tuningを採用し、ポリシーの隠れ状態を事前学習済みworld encoderから得た未来フレーム表現に整列。 - 未来生成を回避し、訓練負担と推論遅延を削減。

2. 先行研究と比べてどこがすごい?

- 従来のWorld-Modeling VLAは明示的な未来生成に依存し、訓練負担と推論遅延が生じる。 - RAF-VLAは未来生成を必要とせず、これらの制限を回避。 - NAVSIMベンチマークで、最先端のVLAプランナーと競合する計画性能を、はるかに少ない訓練サンプル数で達成。 - 訓練オーバーヘッドはわずか3.8%、推論オーバーヘッドは無視できる1 ms。

3. 技術・手法の肝は?

- Future-Aligned Supervised Fine-Tuningを採用。 - 運転行動を学習しながら、ポリシーの隠れ状態を事前学習済みworld encoderから得た未来フレーム表現に整列させる単純な正則化を適用。 - この整列により、未来生成に関連する訓練負担と推論遅延を回避。 - 未来フレーム表現を直接ガイダンスとして使用し、計画に関連する内部表現を形成。

4. どうやって有効だと検証した?

- NAVSIMベンチマークで広範な実験を実施。 - 最先端のVLAプランナーと比較して、競合する計画性能を達成。 - 訓練サンプル数が大幅に少ないことを確認。 - 訓練オーバーヘッドが3.8%、推論オーバーヘッドが1 msであることを測定。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- World-Modeling VLA(具体的な論文名は要旨に記載なし) - NAVSIMベンチマーク - Vision-Language-Action (VLA) モデル - 事前学習済みworld encoder

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dogun Kim, Yongjae Lee, Joonhee Lim, Yeina Lee, Junhyeok Park, Moogeun Park, Dongsuk Kum

分類: cs.RO

原文アブストラクト

Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.

関連論文

PR本紙発行元 EmplifAI