Stream3Dv2: 幾何学的・意味的融合によるストリーミングゼロショット3Dシーン理解の強化
Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding
本論文は、ストリーミングRGB-D入力とノイズのある2Dセグメンテーションマスクを扱うためのトレーニング不要の3Dシーン理解フレームワークを提案し、幾何学的・意味的融合と多様体距離に基づく点群精緻化により性能を向上させる。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
分類: cs.CV
原文アブストラクト
Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supervised methods. However, deploying these models in real-world scenarios is severely hindered by their inability to efficiently handle streaming RGB-D inputs and their inherent vulnerability to noise 2D segmentation masks. To address these critical limitations, we propose Stream3Dv2, a novel training-free framework designed for robust streaming 3D perception. Stream3Dv2 processes sequential data through an original nested local-to-historical architecture, capturing multi-view consistency while circumventing the high computational overhead so as to support timely responses. At its core, we introduce a comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems. Furthermore, we present an innovative manifold-distance-based point cloud refinement strategy. This approach leverages local manifold graphs for point-to-manifold optimization that mitigates the boundary delineation failures caused by Euclidean-distance metrics, and employs geometric bounding boxes to dynamically activate and update historical instances for achieving rapid manifold-to-manifold refinement. Extensive experiments on public datasets demonstrate that Stream3Dv2 consistently outperforms existing baselines in foundational open-vocabulary streaming 3D segmentation and detection. Finally, we show that integrating our framework with an LLM-based agent enables advanced language-driven 3D scene understanding, underscoring its potential for open-world embodied intelligence. Code will be updated at https://github.com/SubmissionsIn/Stream3D.
関連論文
- GroupForward: インスタンスグループ化フィードフォワードガウシアンスプラッティングによる参照可能な3Dシーン構築3Dシーン理解
- CausalSplat: 3Dガウシアンスプラッティングにおける包括的階層的推論に向けて3Dシーン理解
- SmartMage: 3Dシーン理解のための動的モダリティ編成3Dシーン理解
- GPOcc++: 視覚幾何学事前情報を用いた統合スパースガウス占有予測3Dシーン理解
- 孤立したオブジェクトを超えて:3Dシーングラフ解析による関係認識型オープンボキャブラリシーン理解3Dシーン理解
- 2D検出器を用いた3Dガウシアンのオープンボキャブラリおよび参照セグメンテーション3Dシーン理解