WorldSculpt: 接地動画からの構成的3D世界生成
WorldSculpt: Generating Compositional Worlds from Grounded Videos
数百の物体が密集したシーンを、個々の物体メッシュの集合として構成的に3D再構成する手法を提案。単一物体の生成事前分布を多視点観測に適応させ、シーン全体の学習なしで複雑なシーンを生成できることを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
分類: cs.CV
原文アブストラクト
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.
関連論文
- HiSfM: 足場アンカー型階層再構成によるStructure-from-Motionの曖昧性解消3D再構成
- MV-dVRK: 空間的外科知覚のための多視点ベンチマーク3D再構成
- 大規模再構成モデルを用いた人と物体のインタラクション再構成3D再構成
- PIVOT: 実世界3D再構成における姿勢・内部パラメータ・新視点評価のためのマルチ軌道データセットとテストベッド3D再構成
- OccamView: フレーム予算制約下のアクティブ3Dガウス再構成のためのオブジェクト条件付き視点選択3D再構成
- DerainSplat: スパースな雨天視点からのフィードフォワードによるクリーンな3Dガウススプラッティング3D再構成