CLFTv2: 階層的特徴ピラミッドによる効率的なカメラ-LiDAR融合セマンティックセグメンテーション
CLFTv2: Efficient Camera-LiDAR Fusion for Semantic Segmentation via Hierarchical Feature Pyramids
Swinベースのマルチスケールエンコーダと軽量FPNデコーダでカメラとLiDARを融合し、自動運転のセマンティックセグメンテーションを高効率・高精度に実現した研究。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Toomas Tahves, Mauro Bellone, Raivo Sell
分類: cs.CV, cs.RO
原文アブストラクト
Semantic segmentation for autonomous driving requires reliable detection of vulnerable road users (VRUs) despite heavy class imbalance. We introduce CLFTv2, a hierarchical camera-LiDAR fusion framework replacing global ViT attention with a Swin-based multi-scale encoder and a lightweight FPN-style residual decoder. Operating in the 2D perspective domain, CLFTv2 integrates multi-scale geometric cues through shifted-window attention and per-scale residual fusion, avoiding the computational overhead of query-matching decoders. Across three driving datasets, CLFTv2 consistently improves VRU recall. On ZOD, CLFTv2-Large achieves 53.5\% mIoU, improving pedestrian IoU from 35.5\% to 44.9\% over the prior CLFT model. On Waymo, CLFTv2 reaches 61.7\% mIoU. Additionally, a modality-isolation study suggests ViT's global receptive field yields stronger fusion gains only under dense LiDAR returns. Compared to a Swin-based Mask2Former adaptation, CLFTv2 requires 1.4$\times$ fewer GFLOPs and delivers 2.2$\times$ higher throughput, while achieving comparable overall accuracy. These results demonstrate that hierarchical local-attention fusion offers an efficient, scalable alternative to global-attention and query-based decoders for real-time on-vehicle perception in intelligent transportation systems. Source code is publicly available.
関連論文
- Contextrast++: セマンティックセグメンテーションのためのロバストなマルチスケール文脈対比学習セマンティックセグメンテーション
- ICRA 2026 GOOSE 2D細粒度セマンティックセグメンテーションチャレンジ技術報告:フィールドロボティクスにおける堅牢な屋外シーン理解のためのDINOv3活用セマンティックセグメンテーション
- ICRA 2026 GOOSE 2D細粒度セマンティックセグメンテーションチャレンジ技術報告:屋外シーン理解のためのクエリベースセグメンテーションと空間コンテキスト拡大の探求セマンティックセグメンテーション
- Invascal: 不確実性を考慮したLiDARレンジビューセマンティックセグメンテーションのための逆空孔自己校正セマンティックセグメンテーション
- オフロード環境におけるセマンティックセグメンテーションの分布シフト緩和法セマンティックセグメンテーション
- 道路データセットにおけるSAM拡張セグメンテーション:自動運転における重要クラスのバランス調整セマンティックセグメンテーション