日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2609.34085

AD-E2E-JEPA: エンドツーエンド自動運転のための結合埋め込み予測アーキテクチャ

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

自動運転向けのJEPA世界モデルを評価し、計画パッチを16倍削減するSIGReg正則化プロジェクタを導入して100倍の推論高速化を実現した。

詳しい要約

1. どんなもの?

- 本論文は、end-to-end autonomous driving (E2EAD) のための world model として、action-conditioned joint-embedding predictive architecture (JEPA) を体系的に評価し、新たに AD-E2E-JEPA を提案する。 - 既存 JEPA 系 world model (LeWM, DINO-WM, JEPA-WM) を goal-conditioned zero-shot planning 設定で比較し、精度と計算効率のトレードオフを明らかにする。 - AD-E2E-JEPA は SIGReg-regularized learnable projector を projected patch embeddings に適用し、planning patches を 16×、embedding 次元を 4×削減、推論を 100×高速化する。 - 8-frame rollout で 256 candidate trajectories を 0.8 秒で処理し、driving policy…

2. 先行研究と比べてどこがすごい?

- 既存 JEPA 系 world model は、driving に高精度だが計算コストが高いか、計算効率は良いが planning に不十分かのどちらかであった。 - AD-E2E-JEPA は SIGReg-regularized learnable projector により、このトレードオフを解消し、planning 性能を維持しつつ 100×の推論高速化を実現する。 - 従来は driving policy の学習が必要だったが、本手法は policy 学習なしの zero-shot planning で目標到達を達成する点が異なる。 - 自己教師あり事前学習した projector が下流の imitation learning 性能を 80.2 から 85.4 EPDMS に改善することを示す。

3. 技術・手法の肝は?

- action-conditioned JEPA world model を E2EAD に適用し、goal-conditioned zero-shot planning で評価する。 - AD-E2E-JEPA は projected patch embeddings に SIGReg-regularized learnable projector を導入する。 - projector により planning patches を 16×、embedding 次元を 4×削減し、推論を 100×高速化する。 - 8-frame rollout で 256 candidate trajectories を 0.8 秒で処理する。 - world-model rollouts を trajectory vocabularies (256/8,192 candidates) 上で行い、目標へ到達する。

4. どうやって有効だと検証した?

- goal-conditioned zero-shot planning 設定で、ground-truth future observations を goals として使用し、driving policy を学習せずに評価する。 - 平均 20 m 先の目標に対し、256/8,192 candidates の trajectory vocabularies でそれぞれ 4.0/2.8 m の変位で到達することを示す。 - NAVSIMv2 benchmark で、multiplicative safety metrics ありで 67.3/72.9 EPDMS、なしで 84.1/86.5 EPDMS† を達成する。 - 自己教師あり事前学習した projector が下流 imitation learning を 80.2 から 85.4 EPDMS に改善することを実験で示す。

5. 議論はある?

- 既存 JEPA 系 world model の精度と計算効率のトレードオフを指摘し、AD-E2E-JEPA がそれを緩和することを議論する。 - policy 学習なしの zero-shot planning で world model 単体の planning 能力を評価する設定の意義を述べる。 - 自己教師あり事前学習 projector の下流 imitation learning への有効性を議論する。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- LeWM - DINO-WM - JEPA-WM - NAVSIMv2 benchmark に関連する研究 - joint-embedding predictive architecture (JEPA) の原論文

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

関連論文

PR本紙発行元 EmplifAI