MultiFly: 実世界マルチモーダル航空データセットと注釈効率の良いラベル転送およびクロスモーダル意味整合性
MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency
RGB・熱・LiDAR・レーダーの4モダリティを同期した低高度UAVデータセットを構築し、少数の手動注釈から幾何表現を介してラベルを転送することで、大規模な意味セグメンテーション用データを効率的に生成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Markus Gross, Andreas Greiner, Taehyoung Kim, Sivasubiramaniam Subbiah, Tomaž Cotič, Sai Bharadwaj Matha, Conrad Christoph, Oussema Dhaouadi, Simon Zieher, Surya Vijaya Kumar, Gordon Elger, Henri Meeß, Olaf Wysocki, Paul Spannaus, Daniel Cremers
分類: cs.RO, cs.CV
原文アブストラクト
We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly provides 17,272 synchronized samples from four suburban scenes with frame-wise annotations for 15 semantic classes, together with calibration and GNSS-RTK/IMU measurements. To avoid costly and inconsistent modality-specific annotation, we propagate labels from only 115 manually annotated RGB images through shared geometric representations to all four modalities. This approach generates semantic labels for 17,157 additional RGB images, 17,272 thermal images, 840M LiDAR points, and 3.4M radar points. Transferred annotations achieve 89.93% average agreement with held-out manual annotations, and 90.94% average semantic consistency across all six modality pairs. We further establish semantic segmentation benchmarks for all four modalities, revealing distinct architectural behavior for dense LiDAR and sparse radar data. Taken together, MultiFly provides a scalable foundation for multimodal aerial perception and, to the best of our knowledge, the first public real-world low-altitude aerial benchmark that combines consistent frame-wise semantic annotations for RGB, thermal, LiDAR, and radar. Data at https://github.com/markus-42/multifly.
関連論文
- GlassFormer: レーダーと深度の融合によるリアルタイムガラスセグメンテーションの学習マルチモーダル認識
- 対数尤度比融合によるドローン・移動ロボット遠隔操作のための解釈可能なマルチモーダルジェスチャ認識マルチモーダル認識
- 照明条件に応じてRGBと赤外線を適応融合する自律追尾タレットシステムの評価マルチモーダル認識
- OmniUnet: RGB・深度・熱画像を用いた惑星探査ローバー向け非構造地形セグメンテーションのマルチモーダルネットワークマルチモーダル認識
- 敵対的特徴分離による頑健な手術ワークフロー認識のためのマルチモーダルグラフ表現学習マルチモーダル認識
- 海上マルチシーン認識のための軽量マルチモーダルAIフレームワークマルチモーダル認識