PixGS: 画素空間拡散による直接3Dガウシアンスプラット生成
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
3Dガウシアンスプラットを直接生成する単一ステージのパイプラインを提案し、損失のある潜在空間圧縮を回避しつつ、2D生成事前知識を活用して高品質な3Dアセットを高速生成する。
著者: Cao Duy, Phong Nguyen-Ha
分類: cs.CV, cs.GR, cs.RO
原文アブストラクト
Recent advances in 3D content generation from text or images have achieved impressive results, yet view inconsistency from 2D generators and the scarcity of high-quality 3D data remain significant bottlenecks. Existing solutions typically adapt large-scale pre-trained text-to-image latent diffusion models to generate 3D Gaussian Splats (3DGS). However, these approaches often rely on training complex cascade pipelines that are computationally expensive and scalability-limited. Most critically, the quality of generated 3D assets is inherently constrained by each component capacity and compressed latent space, leading to decoding artifacts and accumulated errors. To address these limitations, we propose PixGS, a single-stage pipeline for direct high-quality 3DGS generation, which leverages recent advances in pixel-space diffusion to bypass lossy latent compression while still benefiting from the vast 2D generative priors. By directly denoising 3D Gaussian attributes at each timestep, our method enables precise, splat-level regularization of both appearance and geometry. Furthermore, we introduce a comprehensive supervision strategy that incorporates surface normals, depth, and high-frequency structural information, which is often overlooked in prior works. Experiments demonstrate that PixGS outperforms current state-of-the-art methods while maintaining a fast inference speed (1s on a single A100 GPU), offering a robust and efficient alternative to multi-stage generation pipelines.