日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.27747

言語より潜在表現:運転向けアノテーション効率の良いVLA

Less Language, More Latents: Annotation-Efficient VLAs for Driving

シェア:XThreadsFacebookLINEはてブBluesky

言語注釈が少ない運転データでも、潜在アクションモデルと言語翻訳器を組み合わせてVLAを訓練し、5%未満の言語注釈で完全教師ありに匹敵する性能を達成した。

詳しい要約

1. どんなもの?

- 自動運転向けのVision-Language-Actionモデル(VLA)の訓練において、言語アノテーションの不足を解決する手法LADAを提案。 - LADAは、未ラベルの観測-軌道ペアを活用し、言語条件付き制御を可能にする3段階のパイプライン。 - 第1段階:ベクトル量子化ボトルネックを持つ潜在行動モデルを訓練し、高レベル車両意図のコードブックを生成。 - 第2段階:少数の言語アノテーション付きサブセットを用いて、視覚言語翻訳器を訓練し、観測と言語指示をコードブックにマッピング。 - 第3段階:完全な未ラベルコーパス上で観測-潜在行動ペアを用いて運転VLAを訓練。

2. 先行研究と比べてどこがすごい?

- 従来のVLA訓練は、自然言語指示とペアになったフレームの不足がボトルネック。 - LADAは、言語アノテーションを5%未満に抑え、補助的なchain-of-thought推論やvisual question answeringストリームを使用せずに、完全教師ありベースラインに匹敵または上回る性能を達成。 - 具体的には、閉ループBench2DriveベンチマークでDriving Score 87.98、Success Rate 70.46%を記録。

3. 技術・手法の肝は?

- 3段階のパイプライン: - 潜在行動モデル:ベクトル量子化ボトルネックを用いて、高レベル車両意図のコンパクトなコードブックを学習。 - 視覚言語翻訳器:少数の言語アノテーション付きデータで、観測と言語指示をコードブックにマッピングするように訓練。 - 運転VLA:未ラベルの観測-潜在行動ペアで訓練。 - 言語アノテーションを直接使わず、潜在行動を介することでアノテーション効率を向上。

4. どうやって有効だと検証した?

- 閉ループBench2Driveベンチマークで評価。 - 言語アノテーションを5%未満使用し、Driving Score 87.98、Success Rate 70.46%を達成。 - 完全教師ありベースラインと比較して、同等以上の性能を示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:完全教師ありベースライン、Bench2Driveベンチマーク。 - 関連手法:Vision-Language-Actionモデル(VLA)、潜在行動モデル、ベクトル量子化、チェーンオブソート推論、visual question answering。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania

分類: cs.RO, cs.LG

原文アブストラクト

Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.

関連論文

PR本紙発行元 EmplifAI