GlassFormer: レーダーと深度の融合によるリアルタイムガラスセグメンテーションの学習
GlassFormer: Learning Real-time Glass Segmentation using Radar-Depth Fusion
ミリ波レーダーとRGB-Dを融合し、視覚や深度が苦手な透明ガラス面をリアルタイムにセグメンテーションする軽量トランスフォーマーネットワークを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
分類: cs.CV, cs.RO
原文アブストラクト
Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.
関連論文
- 対数尤度比融合によるドローン・移動ロボット遠隔操作のための解釈可能なマルチモーダルジェスチャ認識マルチモーダル認識
- 照明条件に応じてRGBと赤外線を適応融合する自律追尾タレットシステムの評価マルチモーダル認識
- OmniUnet: RGB・深度・熱画像を用いた惑星探査ローバー向け非構造地形セグメンテーションのマルチモーダルネットワークマルチモーダル認識
- 敵対的特徴分離による頑健な手術ワークフロー認識のためのマルチモーダルグラフ表現学習マルチモーダル認識
- 海上マルチシーン認識のための軽量マルチモーダルAIフレームワークマルチモーダル認識
- 風車ブレードの損傷検出における熱画像とRGB画像の統合活用マルチモーダル認識