日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D生成arXiv:2609.25741

Fysiverse-3D-Vision: 画像から実行可能な3D世界を生成する統合空間推論フレームワーク

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

単一画像から物体の意味・形状・配置を統合的に推論し、物理シミュレーション可能な3Dシーンを生成する視覚言語幾何フレームワークを提案。

詳しい要約

1. どんなもの?

- 単一画像から実行可能な3Dシーンを生成する統合vision-language-geometryフレームワーク - 空間推論と幾何再構成を相互強化する共有表現を確立 - オブジェクトレイアウトを個別のasset generatorの制約を超えて推論可能 - テキスト監督、意味的視覚手がかり、幾何表現を統合Transformerで統合 - オブジェクト条件付きレイアウトモジュールがtranslation, rotation, scaleを予測 - インタラクティブ編集、物理シミュレーション、embodied applicationsに適応

2. 先行研究と比べてどこがすごい?

- 既存の3D生成手法は視覚的に妥当な物体・シーンを合成できるが、空間レイアウト推定が特定のasset generatorに結合 - オブジェクト意味論、メトリック幾何、シーン級空間関係を共同モデル化するのが困難 - 提案手法は空間レイアウト推論とasset synthesisを分離し、適応可能なインターフェースを提供 - 幾何一貫性、レイアウト推定、レンダリング品質、物理特性理解で既存手法を上回る

3. 技術・手法の肝は?

- 統合Transformerでテキスト監督、意味的視覚手がかり、幾何表現を統合 - オブジェクト条件付きレイアウトモジュールがcross-attentionで対象物体表現とグローバル幾何特徴を処理 - 物体のtranslation, rotation, scaleを予測 - 訓練は幾何-言語アラインメントを漸進的に学習 - 再構成能力を保持しつつレイアウト推論を導入 - collision-aware optimizationで物理的一貫性を洗練

4. どうやって有効だと検証した?

- 実験により、既存手法と比較して幾何一貫性、レイアウト推定、レンダリング品質、物理特性理解で優位性を実証 - 具体的なデータセットや評価指標は要旨からは不明

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない - 同分野の定番として、image-conditioned 3D generation, vision-language-geometry, spatial reasoning, executable 3D scene generation に関する論文を挙げる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang

分類: cs.CV

原文アブストラクト

Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset generators. They struggle to jointly model object semantics, metric geometry, and scene-level spatial relationships, which are essential for interactive editing, physical simulation, and embodied applications. We propose Fysiverse-3D-Vision, a unified vision-language-geometry framework for generative 3D scene reconstruction and executable asset construction from a single image. We establish a shared representation where spatial reasoning and geometric reconstruction mutually enhance each other, allowing object layouts to be inferred beyond the constraints of individual asset generators. Our model integrates textual supervision, semantic visual cues, and geometric representations within a unified Transformer to capture scene context, metric geometry, and object-level interactions. An object-conditioned layout module performs cross-attention between target object representations and global geometric features to predict object translation, rotation, and scale. Training progressively learns geometry-language alignment, introduces layout reasoning while preserving reconstruction capability, and refines physical consistency through collision-aware optimization. By separating spatial layout reasoning from asset synthesis, Fysiverse-3D-Vision provides an adaptable interface for interactive scene editing, object-level manipulations, and executable 3D content generation. Experiments demonstrate that our framework achieves superior geometric consistency, layout estimation, rendering quality, and physical property understanding compared with existing approaches.

関連論文

PR本紙発行元 EmplifAI