日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2609.30436

WALT: 自動運転のための世界モデル整合潜在軌道学習

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

凍結した運転世界モデルの知識を軌道潜在空間に転移し、生のウェイポイント生成より高精度な計画を実現する手法を提案。

詳しい要約

1. どんなもの?

- 自動運転のための軌道計画手法 WALT を提案。 - 視覚的世界モデルと生の幾何学的軌道のミスマッチを解消する。 - 凍結した pretrained driving world model から情報を転送し、コンパクトな生成軌道潜在空間を学習。 - 生の waypoints を直接生成せず、dual-branch trajectory autoencoder でマッピング。 - NAVSIM ベンチマークで評価。

2. 先行研究と比べてどこがすごい?

- 従来の driving world models は視覚予測精度が高くても計画に有効とは限らない。 - 生の waypoints を直接生成する baseline に対し、WALT は世界モデルの表現を保持しつつ行動関連情報を抽出。 - 軌道計画の FLOPs を 30.5% 削減。 - PDMS を 89.4 から 89.8 に、EPDMS を 87.3 から 87.9 に改善。

3. 技術・手法の肝は?

- 凍結した pretrained driving world model を変更せずに情報転送。 - dual-branch trajectory autoencoder で waypoints をコンパクト表現に変換。 - 視覚世界モデルから軌道空間へ意味知識を転送。 - 未来の運動と計画に関連するシーン cues を捉える行動表現を学習。 - JEPA と REPA に基づく潜在学習と特徴アラインメントを系統的に検討。

4. どうやって有効だと検証した?

- NAVSIM ベンチマークで評価。 - NAVSIMv1 で PDMS が 89.4 から 89.8 に向上。 - NAVSIMv2 で EPDMS が 87.3 から 87.9 に向上。 - 軌道計画の FLOPs を 30.5% 削減。

5. 議論はある?

- 視覚世界表現を保持しつつ行動関連情報を抽出することが、世界モデルベース軌道計画の有効なインターフェースとなることを示唆。 - 軌道のみの表現学習が下流の計画に与える影響を調査。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- Joint-Embedding Predictive Architectures (JEPA) - Representation Alignment (REPA) - NAVSIM benchmarks - driving world models

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin

分類: cs.RO, cs.CV

原文アブストラクト

Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.

関連論文

PR本紙発行元 EmplifAI