日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.35311

RoGSW4RLD: ロボット世界モデルのロールアウトのためのフィードフォワード4Dガウシアンリフティング

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

シェア:XThreadsFacebookLINEはてブBluesky

複数カメラの映像予測を、視点と時間をまたいで問い合わせ可能な統一的な4Dガウシアン場に変換するフィードフォワード手法を提案し、新規視点の画質・深度・ロボット変位精度を大幅に改善した。

詳しい要約

1. どんなもの?

- 複数カメラの action-conditioned video world model の出力を、統一された time-queryable な metric 4D Gaussian field に持ち上げる feed-forward framework。 - 既存 world model が生成する visual future を直接再構成し、別途 geometric transition model を学習しない。 - 対象は移動する robot-mounted cameras と固定 external views の同期 multi-camera rollouts。

2. 先行研究と比べてどこがすごい?

- 既存の 4D reconstruction 手法は各 camera stream を独立に再構成・統合するため cross-view consistency を強制できない。 - 特に移動する robot-mounted cameras と固定 external views の統合で問題が顕著。 - RoGSW4RLD は camera-wise reconstruction with calibrated merging を大きく上回る。

3. 技術・手法の肝は?

- 2段階アーキテクチャ。 - Stage 1: cross-view evidence と robot-specific articulated geometry・kinematics を融合し metric 4D field を jointly 形成。 - Stage 2: 初期の temporal displacements を厳密に保持しつつ geometry と appearance を refine。

4. どうやって有効だと検証した?

- 256 held-out DROID episodes で評価。 - camera-wise reconstruction with calibrated merging に対し novel-view PSNR を 2.15 dB 改善、depth AbsRel を 47% 削減、robot displacement error を 61% 低減。 - action-conditioned Cosmos 3 rollouts でも同様の gains を確認。

5. 議論はある?

- 予測された video futures を consistent で spatially queryable な 4D metric representations に変換できることを示す。 - 限界や失敗ケース、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- DROID - Cosmos 3 - action-conditioned video world models - 4D reconstruction methods - 4D Gaussian field 関連手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim

分類: cs.CV, cs.RO

原文アブストラクト

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.

関連論文

PR本紙発行元 EmplifAI