Lang3DSeg: 点群トランスフォーマによるアノテーションフリーなオープンボキャブラリ3Dセグメンテーション
Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers
2D視覚言語モデルの出力をLiDARに投影して学習する、アノテーション不要のオープンボキャブラリ3Dセグメンテーション手法。点群トランスフォーマを屋外LiDARのバックボーンとして初めて採用し、投影ラベルの深度曖昧性をクラス優先度と深度ギャップで補正する。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li
分類: cs.CV, cs.RO
原文アブストラクト
Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.
関連論文
- SplatLabel: 4Dガウシアンスプラッティングによる擬似ラベリング3Dセグメンテーション
- UniPart: 実世界インタラクションのためのゼロショット言語接地3Dパーツセグメンテーション3Dセグメンテーション
- 仮想ドローンを飛ばして3Dガウシアンをオンラインでセグメンテーション3Dセグメンテーション
- EPS3D: エンドツーエンドのフィードフォワード型3Dパノプティックセグメンテーション3Dセグメンテーション
- T-FunS3D: タスク駆動型階層的オープンボキャブラリ3D機能セグメンテーション3Dセグメンテーション
- TrackRef3D: 3Dガウススプラッティングにおけるオープンワールド参照セグメンテーションのための多視点一貫トラック・アンド・ラベル手法3Dセグメンテーション