日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18623

FIVE-VLA: 再帰的行動メモリによる高速で効果的な自動運転

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

シェア:XThreadsFacebookLINEはてブBluesky

高解像度画像をわずか98トークンに圧縮する効率的な視覚エンコーダと、過去の行動トークンを条件に用いる再帰的行動メモリ(RAM)を組み合わせ、641Mパラメータで従来のVLAより高速・高精度に軌道予測を行う自動運転モデル。

詳しい要約

1. どんなもの?

- 自動運転向けの Vision-Language-Action model (VLA) である FIVE-VLA を提案 - パラメータ数 641M の軽量モデル - 高解像度画像 (448×896) を 98 token のみで処理する効率的 vision encoder を採用 - テキスト生成を完全に省略し、単一パスで trajectory 予測を実行 - Recurrent Action Memory (RAM) により過去の action token を条件付け、時間的文脈を付与 - 閉ループ運転ベンチマーク Bench2Drive や大規模実世界データセット NVIDIA Physical AI AV dataset で評価

2. 先行研究と比べてどこがすごい?

- 従来の VLA はパラメータ数が過大、高解像度画像処理が非効率、時間的記憶が欠如 - 既存手法より 5 倍以上少ない 98 token で高解像度画像を処理 - テキスト生成を省略し単一パスで trajectory を予測 - Bench2Drive 閉ループベンチマークで、先行 SOTA VLA より交通規則違反なしで約 10% 多くのルートを完走 - NVIDIA Physical AI AV dataset の非反応オープンループシミュレーションで、SimLingo より衝突違反率が single-view で 10.2%、four-view で 7.7% 低い - A100 で約 30 fps、T4 で約 4 fps 動作し、従来手法より 8〜30 倍高速

3. 技術・手法の肝は?

- 効率的 vision encoder により高解像度画像 (448×896) を 98 token に圧縮 - テキスト生成をバイパスし、単一パスで trajectory を直接予測 - Recurrent Action Memory (RAM) を導入:軽量モジュールで、過去の action token を条件として action 予測を行う - RAM により overtaking や emergency braking などの manoeuvre に必要な時間的文脈を提供 - 全体で 641M パラメータの VLA アーキテクチャ

4. どうやって有効だと検証した?

- 閉ループ運転ベンチマーク Bench2Drive で評価 - 大規模実世界データセット NVIDIA Physical AI AV dataset を用いた非反応オープンループシミュレーション - 先行 SOTA VLA と比較して交通規則違反なしのルート完走率が約 10% 向上 - SimLingo と比較して衝突違反率が single-view で 10.2%、four-view で 7.7% 低減 - 推論速度を A100 で約 30 fps、T4 で約 4 fps と測定し、従来手法比 8〜30 倍の高速化を確認

5. 議論はある?

- パラメータ数削減、高解像度画像処理の効率化、時間的記憶の欠如という VLA の課題に対処 - RAM が overtaking や emergency braking などの manoeuvre に有効であることを示唆 - ただし、RAM の詳細な設計や限界、他の manoeuvre への影響については要旨からは不明 - オープンループ評価は非反応シミュレーションであり、実世界での反応性や安全性に関する議論は要旨からは不明 - エッジデバイス (T4) での動作は示されているが、実際の車載環境での性能や消費電力については要旨からは不明

6. 次に読むべき論文は?

- SimLingo (比較対象として言及) - 先行 state-of-the-art VLA (具体的名称は要旨からは不明) - Bench2Drive (ベンチマーク) - NVIDIA Physical AI AV dataset (データセット) - 関連手法として Vision-Language-Action models (VLA) 全般 - 時間的記憶を扱う Recurrent Action Memory (RAM) に関連する研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao, Puneet K. Dokania

分類: cs.CV, cs.RO

原文アブストラクト

State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.

関連論文

PR本紙発行元 EmplifAI