日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2610.08780

DepthWorld: ロボットマニピュレーションのための3Dワールドモデル

DepthWorld: 3D World Model for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作向けに、多視点RGBと深度を同時予測する3DワールドモデルDepthWorldを提案し、DROIDデータセットをキャリブレーションしたDROID-3Dで学習することで、RGB予測精度の向上と正確なメートル深度推定を実現した。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションのための3Dワールドモデル「DepthWorld」を提案。 - ビデオベースのワールドモデルはRGBのみで訓練され、フレーム単位では正しく見えても一貫した3D世界を構成できない問題に対処。 - 大規模3D監督と、事前学習済みビデオ事前分布を乱さないアーキテクチャの両立を目指す。 - DROIDデータセットにキャリブレーションパイプラインを適用し、DROID-3Dを構築。 - Stable Video Diffusionをベースに、マルチビューRGBと深度を同時予測するモデルを訓練。

2. 先行研究と比べてどこがすごい?

- 従来のビデオベースワールドモデルはRGBのみで訓練され、3D幾何一貫性に欠ける。 - 本研究は大規模3D監督(DROID-3D)と、事前学習済みVAEを変更せずに深度を組み込むアーキテクチャを提供。 - 同一のRGBのみのベースラインと比較して、RGB予測自体が+1.48 dB PSNR改善。 - 同時に正確なメートル深度を生成し、下流の幾何推論を可能にする。

3. 技術・手法の肝は?

- キャリブレーションパイプライン:学習済みステレオ深度とジョイントファクターグラフを組み合わせ、同一物理ロボットから収集した全エピソードをプールして共有キネマティックパラメータとシーンごとの外部パラメータを復元。 - DROIDデータセットに適用し、DROID-3Dを構築(高密度メートル深度と再キャリブレーションされたマルチビュー外部パラメータを提供、外部カメラで90%のエピソードが<0.7 px再投影誤差)。 - DepthWorld:Stable Video Diffusionベースのワールドモデルで、空間潜在タイリングによりマルチビューRGBと深度を同時予測。事前学習済みVAEは変更しない。

4. どうやって有効だと検証した?

- DROIDデータセットにキャリブレーションパイプラインを適用し、DROID-3Dを構築。外部カメラで90%のエピソードが<0.7 px再投影誤差を達成。 - DepthWorldを訓練し、同一のRGBのみのベースラインと比較してRGB予測が+1.48 dB PSNR改善。 - 同時に正確なメートル深度を生成し、下流の幾何推論に利用可能であることを示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Stable Video Diffusion(ベースモデル) - DROID dataset(使用データセット) - 関連手法:ビデオベースワールドモデル、ステレオ深度推定、ジョイントファクターグラフ最適化

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jai Bardhan, Josef Sivic, Vladimir Petrik

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.

関連論文

PR本紙発行元 EmplifAI