日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D生成arXiv:2607.16805v1

Scene-SAM3D: 微調整不要の多視点シーン資産生成

Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning

シェア:XThreadsFacebookLINEはてブBluesky

単一画像から3Dオブジェクトを生成するSAM3Dを、キャリブレーション済みの多視点画像から3Dシーン資産を生成するフレームワークに拡張した。冗長な視点を削減しつつ、隠れた領域の情報を補完し、潜在速度融合と軽量なガウス最適化により高品質なシーンを生成する。

著者: Yuqi Zhang, Yadan Luo, Xiangyu Sun, Fengyi Zhang, Zi Huang, Xin Tan

分類: cs.CV

原文アブストラクト

High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-image 3D generation models such as SAM3D remain insufficient for real-world scenes, where severe occlusions, redundant observations, and cross-view inconsistencies make reliable scene generation challenging. We introduce Scene-SAM3D, a training-free framework that extends SAM3D from single-view object generation to calibrated multi-view scene asset generation. Scene-SAM3D selects a compact set of complementary views, reducing observation redundancy while providing additional evidence for regions occluded in individual views. Based on the selected views, it performs step-efficient latent velocity fusion to integrate multi-view evidence and suppress cross-view conflicts in canonical space. Finally, a lightweight rigid-object Gaussian optimization refines the scene layout within 200 iterations while preserving the generated object geometry. Experiments on Replica and ScanNet++ demonstrate consistent improvements at both instance and scene levels, with our method reducing scene-level CD by 43.8% on Replica and 30.9% on ScanNet++, while cutting flow-model sampling FLOPs and wall-time latency by nearly 20% under the same multi-view setting. Code will be released at https://github.com/xibi777/Scene-SAM3D.

関連論文