Spheriverse: 実世界の球面観測による3Dシーン理解
Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild
球面画像とLiDARの大規模データセットを構築し、球面幾何を考慮した占有予測フレームワークSphereOccを提案して、セマンティック占有予測や3D物体検出のベンチマークで優れた性能を示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Fei Teng, Sheng Wu, Mengfei Duan, Guoqiang Zhao, Junhui Ma, Kai Luo, Siyu Li, Hao Shi, Zhiyong Li, Kailun Yang
分類: cs.CV, cs.RO, eess.IV
原文アブストラクト
Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates. This cross-space representation gap complicates geometric correspondence and semantic evidence aggregation. To delve into this challenge, we introduce Spheriverse, comprising $64,400$ temporally aligned spherical image-LiDAR pairs organized into 644 sequences. The dataset spans diverse scenes, illumination, and weather conditions, with fine-grained semantic classes. We further establish benchmarks for semantic occupancy prediction, semantic mapping, and 3D object detection, evaluating 30+ methods through overall and scene-wise comparisons. For dense prediction, we propose SphereOcc, an occupancy framework that couples spherical geometry modeling with semantic evidence retrieval. Cartesian-Spherical Representation Remodeling (CSRR) incorporates spherical range-azimuth geometry into Cartesian voxel features through region-wise modulation. Spherical Evidence Re-querying (SER) then conditions queries on voxel content and range-height-azimuth geometry to adaptively retrieve relevant semantic evidence from source spherical image features. SphereOcc achieves 13.91% mIoU and 24.65% GeoIoU, outperforming the respective best-performing methods, TPVFormer and SurroundOcc, by 1.70 and 2.10 percentage points. It also ranks first in both metrics across all five scenes, with consistent advantages across the evaluated spatial partitions and reduced fields of view. The established benchmark and source code will be available at https://feit-feiteng.github.io/Spheriverse.
関連論文
- Stream3Dv2: 幾何学的・意味的融合によるストリーミングゼロショット3Dシーン理解の強化3Dシーン理解
- GroupForward: インスタンスグループ化フィードフォワードガウシアンスプラッティングによる参照可能な3Dシーン構築3Dシーン理解
- CausalSplat: 3Dガウシアンスプラッティングにおける包括的階層的推論に向けて3Dシーン理解
- SmartMage: 3Dシーン理解のための動的モダリティ編成3Dシーン理解
- GPOcc++: 視覚幾何学事前情報を用いた統合スパースガウス占有予測3Dシーン理解
- 孤立したオブジェクトを超えて:3Dシーングラフ解析による関係認識型オープンボキャブラリシーン理解3Dシーン理解