GPOcc++: 視覚幾何学事前情報を用いた統合スパースガウス占有予測
GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors
視覚幾何学の事前情報を占有予測に適したスパースガウス表現に変換し、多視点・時系列を統一的に扱う3Dシーン理解手法GPOcc++を提案した。屋内・屋外のベンチマークで高い性能と効率性を示した。
著者: Changqing Zhou, Yueru Luo, Yulan Guo, Bing Wang, Jie Qin, Changhao Chen
分類: cs.CV
原文アブストラクト
Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations remains challenging, particularly in occluded and unobserved regions. Visual geometry priors offer strong and generalizable geometric cues for addressing this challenge, but their outputs are inherently surface-centric, whereas occupancy prediction requires reasoning about volumetric interiors and free space. To bridge this gap, we introduce GPOcc, which transforms visual geometry priors into occupancy-aware sparse Gaussian representations for efficient and expressive volumetric scene modeling. Building on GPOcc, GPOcc++ models multi-view observations and temporal sequences within a unified framework, allowing spatial and temporal evidence to be handled through the same representation. We further extend GPOcc++ from indoor scenes to outdoor occupancy prediction. Extensive experiments on both indoor and outdoor benchmarks demonstrate consistently strong performance across both multi-view and temporal settings, together with favorable efficiency and generalization. Code will be released at https://github.com/JuIvyy/GPOcc.
関連論文
- Stream3Dv2: 幾何学的・意味的融合によるストリーミングゼロショット3Dシーン理解の強化3Dシーン理解
- GroupForward: インスタンスグループ化フィードフォワードガウシアンスプラッティングによる参照可能な3Dシーン構築3Dシーン理解
- CausalSplat: 3Dガウシアンスプラッティングにおける包括的階層的推論に向けて3Dシーン理解
- SmartMage: 3Dシーン理解のための動的モダリティ編成3Dシーン理解
- 孤立したオブジェクトを超えて:3Dシーングラフ解析による関係認識型オープンボキャブラリシーン理解3Dシーン理解
- 2D検出器を用いた3Dガウシアンのオープンボキャブラリおよび参照セグメンテーション3Dシーン理解