日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.07381

GeoWM: 明示的幾何における効率的な直接世界モデリング

GeoWM: Efficient Direct World Modeling in Explicit Geometry

シェア:XThreadsFacebookLINEはてブBluesky

観測RGBフレームから幾何基盤モデルで幾何履歴を作り、フローマッチングTransformerで将来のシーン幾何を直接予測する世界モデルを提案。再帰的ロールアウト不要で長期的な予測精度と推論速度を改善。

詳しい要約

1. どんなもの?

- 3D scene geometryとその時間的進化をモデル化するためのworld model - 観測RGBフレームからgeometry foundation modelでgeometric historyを生成 - flow-matching transformerで指定未来horizonのscene geometryを直接予測 - 再帰的rollout不要で、depth, camera pose, 3D scene geometryを予測 - 自動運転、ロボティクス、航空飛行、動的マニピュレーションに適用可能

2. 先行研究と比べてどこがすごい?

- 従来のworld modelは未来画像やlatent表現を予測し、そこからgeometryを復元 - 幾何構造を明示的にモデル化せず、再帰的rolloutで長horizon予測を行うため誤差蓄積と計算コスト増大 - GeoWMは幾何を直接予測し、再帰的rolloutを排除 - 長horizonでの推論時間を大幅削減しつつ、depth, camera pose, 3D scene geometryの予測精度で既存world modelを上回る

3. 技術・手法の肝は?

- geometry foundation modelで観測RGBフレームをgeometric historyに変換 - そのhistoryを条件としてflow-matching transformerが指定未来horizonのscene geometryを予測 - lightweight camera-motion predictorで未来視点を推定 - 観測geometryを予測視点に投影し、未来geometry予測の有効なgeometric priorとして利用

4. どうやって有効だと検証した?

- 都市運転、航空飛行、動的マニピュレーションを含む4つのデータセットで実験 - 未来のdepth, camera pose, 3D scene geometryの予測性能を評価 - 評価対象のworld modelと比較し、GeoWMが優位であることを確認 - 長horizonでの推論時間が大幅に短縮されることを示した

5. 議論はある?

- 再帰的rolloutに依存しない直接予測により誤差蓄積と計算コストを低減 - 明示的幾何モデリングの有効性を示唆 - ただし、要旨からは失敗事例や限界、一般化可能性に関する議論は不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてworld model, flow-matching transformer, geometry foundation model, camera-motion predictorが挙げられる - 同分野の定番としてDreamer, PlaNet, VideoGPT, Geometry-aware world modelsなどが考えられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mehrdad Noori, Guile Wu, Sam Hosseini, Dongfeng Bai

分類: cs.CV

原文アブストラクト

Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.

関連論文

PR本紙発行元 EmplifAI