日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
自動運転arXiv:2610.11161

VGGTWorld-VLA:自動運転のための意図条件付き3D世界進化

VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

VGGTベースの3D世界モデルに運転意図と行動を条件として注入し、同一シーンから異なる未来の3D形状を予測できるようにした自動運転向け手法。

詳しい要約

1. どんなもの?

- VGGTWorld-VLAは、自動運転向けの意図条件付き3D世界進化モデル。 - VGGTを基盤とし、時間的3D予測を拡張。 - 運転意図や行動に条件付けられた未来の3Dシーン幾何を予測。 - 同一観測シーンでも異なるego actionに応じた未来幾何を生成可能。

2. 先行研究と比べてどこがすごい?

- 従来のVGGT拡張は未来進化が運転意図や行動に弱く条件付けられていた。 - 本手法はaction-semantic conditioningを導入し、代替行動依存の未来予測を可能に。 - 幾何・言語・行動を橋渡しするgeometry-language-action bridgeを開発。 - ベースラインと比較して競争力のある幾何予測性能を実証。

3. 技術・手法の肝は?

- action-semantic conditioning mechanism:運転意味論とego-motion表現を未来トークンストリームに注入。 - geometry-language-action bridge:履歴幾何、VLA意味特徴、maneuverとtrajectory表現を適応し、未来幾何予測を共同条件付け。 - これにより同一観測下で異なるego actionに応じた未来幾何を生成。

4. どうやって有効だと検証した?

- NAVSIM上で未来幾何予測を評価。 - 条件付けアブレーションで意味情報と行動情報の寄与を検証。 - ベースラインと比較し競争力のある性能を確認。 - アブレーション結果が意味・行動条件付けの有効性を支持。

5. 議論はある?

- 意味・行動条件付けがVGGTベースの世界予測の制御可能性に寄与する可能性を示唆。 - 自動運転における制御可能な世界予測への応用可能性を議論。 - 具体的な限界や課題については要旨からは不明。

6. 次に読むべき論文は?

- VGGT(基盤モデル) - VGGT-World(時間的3D予測拡張) - NAVSIM(評価データセット) - VLA(Vision-Language-Action)関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhaoyang Liu, Kun Jiang, Ziying Song, Diange Yang

分類: cs.CV, cs.RO

原文アブストラクト

VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations. Although recent extensions enable temporal 3D prediction, their future evolution remains weakly conditioned on driving intentions and actions, limiting their ability to model alternative action-dependent futures. We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving. First, we introduce an action--semantic conditioning mechanism that injects complementary driving semantics and ego-motion representations into the future-token stream, enabling different future geometry predictions for the same observed scene under alternative ego actions. Second, we develop a geometry--language--action bridge that adapts historical geometry, VLA semantic features, and maneuver and trajectory representations for joint conditioning of future geometry prediction. We evaluate future geometry prediction on NAVSIM, while conditioning ablations further examine the contributions of semantic and action information. Compared with the baseline, our method demonstrates competitive geometry prediction performance. Ablation studies further support the effectiveness of semantic and action conditioning. These results demonstrate the potential of semantic and action conditioning for controllable VGGT-based world prediction in autonomous driving.

関連論文

PR本紙発行元 EmplifAI