日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3Dシーン生成arXiv:2608.27073

SpatialCrafter: 生成3Dプロキシを用いた単一画像からのワールドモデリング

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

シェア:XThreadsFacebookLINEはてブBluesky

単一画像から探索可能な3Dシーンを生成する新しい2段階フレームワークを提案し、3Dプロキシ生成と外観精緻化を分離することで、従来のビデオ拡散モデルより高品質で3D整合性の高いシーン生成を実現した。

詳しい要約

1. どんなもの?

SpatialCrafterは、単一の画像から探索可能な3Dシーンを生成するための2段階フレームワークである。従来のvideo diffusion model (VDM)ベースの手法が、スパースな点群や2Dパノラマなどの不完全な条件信号に依存し、確率的な幻覚や長期的なドリフト、3D一貫性の低下を引き起こす問題に対処する。具体的には、グローバルな3Dプロキシを導入し、生成プロセスを「グローバルプロキシ生成」と「外観リファインメント」に分解する。プロキシ生成では、Point-anchored Sparse Structure (PaSS) Flowモジュールを用いて、空間的に整列し幾何学的に一貫した3Dプロキシを予測する。外観リファインメントでは、VDMをGenerative Deferred Refinerとして再構成し、プロキシで定義されたシーン幾何に基づいて高周波のフォトリアリスティックな詳細を合成する。さらに、このタスク用の大規模データセットが存在しないため、115Kシーンからなる新しいハイブリッドデータセットを構築した。

2. 先行研究と比べてどこがすごい?

先行研究のVDMベースの手法は、スパースな点群や2Dパノラマなどの不完全な条件信号に依存しており、これが確率的な幻覚、長期的なドリフト、3D一貫性の低下を引き起こしていた。SpatialCrafterは、グローバルな3Dプロキシを導入することで、これらの問題を解決し、高忠実度の画像からシーン生成を実現している。また、既存のデータセットが存在しないため、新たに115Kシーンのハイブリッドデータセットを構築した点も独自性が高い。

3. 技術・手法の肝は?

手法の肝は、2段階の生成プロセスと、事前学習済みVDMとの統合にある。まず、PaSS Flowモジュールが点群アンカーを用いて3Dプロキシを予測する。次に、Generative Deferred RefinerとしてVDMを再構成し、プロキシの幾何に基づいて外観を合成する。さらに、Parallel Geometry InjectionとProxy-Aware Corruption training戦略を導入し、プロキシのアーティファクトに対するロバスト性を向上させつつ、事前学習済みの生成多様体を乱さないようにしている。

4. どうやって有効だと検証した?

合成データセットと実世界データセットの両方で広範な実験を行い、SpatialCrafterが最先端の手法を上回ることを示した。特に、長期的なドリフトの軽減、高速なカメラモーションや極端な視点変更に対するロバスト性と一貫性を検証している。

5. 議論はある?

要旨からは、議論の余地や限界についての具体的な言及は不明である。ただし、提案手法がプロキシの品質に依存する可能性や、データセットの構築方法に関する詳細な議論が考えられるが、要旨には記載がない。

6. 次に読むべき論文は?

要旨で参照されている関連手法は、video diffusion model (VDM)に基づく既存の画像からシーン生成手法である。具体的な論文名は挙げられていないが、VDMを用いた画像生成やシーン生成に関する研究が関連する。次に読むべき論文としては、VDMを用いた画像生成の基礎論文や、3Dシーン生成における点群やパノラマを用いた手法が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan

分類: cs.CV, cs.RO

原文アブストラクト

Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency. We present SpatialCrafter, a novel two-stage framework that addresses these issues by introducing a global 3D proxy for high-fidelity image-to-scene generation. Specifically, we decompose the generation process into global proxy generation and appearance refinement. For proxy generation, we propose a Point-anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy. For appearance refinement, we re-frame the VDM as a Generative Deferred Refiner which synthesizes high-frequency photorealistic details upon proxy-defined scene geometry. To better integrate the proxy with the pre-trained VDM, we introduce Parallel Geometry Injection and Proxy-Aware Corruption training strategies, which improve robustness to proxy artifacts without disrupting the pretrained generative manifold. Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation. Extensive experiments on both synthetic and real-world datasets show that SpatialCrafter outperforms state-of-the-art methods, mitigates long-term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes. Code, models, and the newly constructed dataset will be publicly released. See more at https://fangchuan.github.io/SpatialCrafter/.

関連論文