日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3DセグメンテーションarXiv:2610.00855

Lang3DSeg: 点群トランスフォーマによるアノテーションフリーなオープンボキャブラリ3Dセグメンテーション

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

シェア:XThreadsFacebookLINEはてブBluesky

2D視覚言語モデルの出力をLiDARに投影して学習する、アノテーション不要のオープンボキャブラリ3Dセグメンテーション手法。点群トランスフォーマを屋外LiDARのバックボーンとして初めて採用し、投影ラベルの深度曖昧性をクラス優先度と深度ギャップで補正する。

詳しい要約

1. どんなもの?

- どんなもの? - 自動運転のための3D LiDARセマンティックセグメンテーション手法。 - アノテーション不要のオープンボキャブラリセグメンテーション。 - point transformerをバックボーンとして使用。 - 屋外環境の3D LiDARに適用。 - 幾何学的事前学習なしでスクラッチから訓練。

2. 先行研究と比べてどこがすごい?

- 先行研究と比べてどこがすごい? - 従来のオープンボキャブラリ手法はvoxelベースのsparse convolutionに依存。 - point transformerは屋内環境に限定されていた。 - 本研究はpoint transformerを屋外3D LiDARに初めて適用。 - アノテーション不要手法の中で最高精度を達成。 - nuScenes validationで52.8% mIoU、SemanticKITTIで41.4% mIoU。

3. 技術・手法の肝は?

- 技術や手法の肝はどこ? - 2D vision-languageモデルの出力をLiDARに投影し3Dネットワークに蒸留。 - 2D-to-3Dラベル投影のノイズに対処。 - 深度曖昧性を解決するため、明示的なクラス優先度ルールでマスクを合成。 - 投影されたインスタンスを深度分布の最初のギャップで切り捨て。 - 投影誤差を直接修正し、登録シーケンスで平均化しない。

4. どうやって有効だと検証した?

- どうやって有効だと検証した? - nuScenes validationとSemanticKITTIで評価。 - アノテーション不要手法の中で最高のmIoUを達成。 - 単一LiDARスイープで3Dセマンティックセグメンテーションを実行。 - 推論はリアルタイムで、vision-languageモデルを実行しない。

5. 議論はある?

- 議論はある? - 要旨からは不明。

6. 次に読むべき論文は?

- 次に読むべき論文は? - 要旨で参照/比較されている研究:voxel-based sparse convolutionsを用いたオープンボキャブラリ手法、point transformerの屋内応用。 - 関連手法:2D vision-languageモデル(例:CLIP)、LiDARセグメンテーション(例:nuScenes, SemanticKITTIベンチマーク)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li

分類: cs.CV, cs.RO

原文アブストラクト

Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.

関連論文

PR本紙発行元 EmplifAI