PhysWAM: 自動運転のための物理的に整合した世界行動モデル
PhysWAM: Physically Consistent World Action Model for Autonomous Driving
マルチビュー動画・深度・自車運動を単一のフローマッチングTransformerで同時生成し、生成した深度と運動をLiDAR点群に整合させる幾何制約CPPを導入した自動運転向け世界行動モデル。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang
分類: cs.RO, cs.CV
原文アブストラクト
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
関連論文
- 自動運転のシーン予測のための拡散トランスフォーマーワールドアクションモデル自動運転/世界モデル
- Discrete-WAM: 世界モデルとポリシー学習のための統一離散ビジョン・アクショントークン編集自動運転/世界モデル
- VectorWorld: ベクトルグラフ上の拡散フローによる効率的なストリーミング世界モデル自動運転/世界モデル
- DynFlowDrive: フローベース動的世界モデルによる自動運転自動運転/世界モデル
- DreamerAD: 潜在世界モデルによる自動運転のための効率的な強化学習自動運転/世界モデル
- UniDWM: 多面的表現学習による統合運転世界モデル自動運転/世界モデル