CausalSplat: 3Dガウシアンスプラッティングにおける包括的階層的推論に向けて
CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting
3Dガウシアンスプラッティングを用いたシーン理解において、暗黙的な意図や常識推論を扱う新しいタスクを提案し、ベンチマークとフレームワークCausalSplatを構築した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jiayu Ding, Meilu Song, Yun Chen, Wei Gao, Ge Li
分類: cs.CV
原文アブストラクト
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat
関連論文
- Stream3Dv2: 幾何学的・意味的融合によるストリーミングゼロショット3Dシーン理解の強化3Dシーン理解
- GroupForward: インスタンスグループ化フィードフォワードガウシアンスプラッティングによる参照可能な3Dシーン構築3Dシーン理解
- SmartMage: 3Dシーン理解のための動的モダリティ編成3Dシーン理解
- GPOcc++: 視覚幾何学事前情報を用いた統合スパースガウス占有予測3Dシーン理解
- 孤立したオブジェクトを超えて:3Dシーングラフ解析による関係認識型オープンボキャブラリシーン理解3Dシーン理解
- 2D検出器を用いた3Dガウシアンのオープンボキャブラリおよび参照セグメンテーション3Dシーン理解