2D検出器を用いた3Dガウシアンのオープンボキャブラリおよび参照セグメンテーション
Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors
3Dガウシアンスプラッティングのシーン表現に、2D物体検出器のセマンティック投票を集約してインスタンス特徴を学習し、言語クエリによるオープンボキャブラリ・参照セグメンテーションを実現する手法を提案した。
著者: Jameel Hassan, Yasiru Ranasinghe, Vishal Patel
分類: cs.CV
原文アブストラクト
3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such as embodied AI. Recent methods achieve this by learning an instance feature attribute and assigning semantics by distilling high-dimensional Contrastive Language-Image Pretraining (CLIP) features directly into the scene representation. However, the instance grouping mechanisms of these methods either require a predefined number of instances or suffer from noise in their bottom-up grouping strategies. Furthermore, the reliance on CLIP restricts semantic understanding to simple noun phrases, preventing complex spatial reasoning and referential expression grounding. We present GaussDet, a method that circumvents the need for dense CLIP features by leveraging discrete, open-vocabulary 2D object detectors with referring expression capabilities. We learn instance features for individual Gaussians to decompose the scene into 3D instance groups. By rendering these groups and aggregating semantic votes from multi-view 2D detections, we generate a robust View-Aggregated Semantic Label Distribution (VASD) for each 3D instance. This view-aggregation strategy acts as a strong regularizer, attenuating spurious labels caused by low-quality instance grouping. Our approach enables a straightforward, zero-shot extension from simple language queries to complex referential grounding. Extensive evaluations across two key tasks -- open-vocabulary segmentation (LeRF-OVS, ScanNet) and referring expression grounding (Ref-LeRF) -- demonstrate that GaussDet achieves consistent improvements over existing methods. Most notably, we achieve a substantial 16.7% mIoU improvement in referential grounding within a strict zero-shot setting.
関連論文
- Stream3Dv2: 幾何学的・意味的融合によるストリーミングゼロショット3Dシーン理解の強化3Dシーン理解
- GroupForward: インスタンスグループ化フィードフォワードガウシアンスプラッティングによる参照可能な3Dシーン構築3Dシーン理解
- CausalSplat: 3Dガウシアンスプラッティングにおける包括的階層的推論に向けて3Dシーン理解
- SmartMage: 3Dシーン理解のための動的モダリティ編成3Dシーン理解
- GPOcc++: 視覚幾何学事前情報を用いた統合スパースガウス占有予測3Dシーン理解
- 孤立したオブジェクトを超えて:3Dシーングラフ解析による関係認識型オープンボキャブラリシーン理解3Dシーン理解