日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転/世界モデルarXiv:2609.37970

PhysWAM: 自動運転のための物理的に整合した世界行動モデル

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

マルチビュー動画・深度・自車運動を単一のフローマッチングTransformerで同時生成し、生成した深度と運動をLiDAR点群に整合させる幾何制約CPPを導入した自動運転向け世界行動モデル。

詳しい要約

1. どんなもの?

- 自動運転向けの統合 world-action model (WAM) である PhysWAM を提案。 - 単一の flow-matching transformer 内で multiview video、metric depth、ego motion を同時に denoise する。 - 世界予測と行動生成を測定された scene geometry に接地させるため Coupled Point Projection (CPP) を導入。 - 推論時の trajectory selection は label-free な consensus rule のみで行う。

2. 先行研究と比べてどこがすごい?

- 従来の WAM は joint generation を行うが、予測間に共有の幾何制約を課していなかった。 - PhysWAM は CPP により生成 depth と ego motion を LiDAR 点群と結びつけ、物理的整合性を直接監督する。 - 学習済み scorer や simulator feedback を使わず、単純な選択規則でも強い planning 性能と zero-shot 転移を達成。 - 要旨からは、比較対象の具体的な先行研究名は不明。

3. 技術・手法の肝は?

- 単一の flow-matching transformer で multiview video、metric depth、ego motion を co-denoise する。 - Coupled Point Projection (CPP): 生成 depth を 3D 点に unproject し、生成 SE(3) ego motion で変換。 - 記録された ego motion で変換した LiDAR 点との距離を最小化する幾何制約を課す。 - この制約は生成 depth と motion を標準の flow-matching 目的と共に jointly supervise する。 - 推論時は label-free consensus rule のみで trajectory を選択。

4. どうやって有効だと検証した?

- NAVSIM v1 および v2 の planning で評価。 - zero-shot closed-loop transfer を検証。 - future video と metric-depth prediction を評価。 - CPP が planning と depth prediction の両方を改善することを示す。 - 単純な選択手順にもかかわらず強い planning 性能と未見環境への zero-shot 転移を確認。

5. 議論はある?

- 要旨からは、限界や議論の詳細は不明。 - 単純な label-free consensus rule でも強い性能が得られる点が強調されている。 - CPP が planning と depth prediction を改善することが示されている。 - 幾何関係が world と action 生成を結合する直接的な手段であると主張。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な研究は不明。 - 同分野の関連手法として world-action models (WAMs)、flow-matching transformer、NAVSIM benchmark が挙げられる。 - 関連研究を探す際は autonomous driving における world model や planning の定番手法を参照するとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang

分類: cs.RO, cs.CV

原文アブストラクト

World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

関連論文

PR本紙発行元 EmplifAI