HAVT-IVD: アイドリング車両検出のための異種モダリティ対応音声視覚ネットワーク
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
監視映像とマルチチャンネル音声を組み合わせ、車両が移動中・アイドリング中・エンジン停止中かを検出するネットワークを提案。視覚特徴ピラミッドと分離ヘッドで異種モダリティ問題に対処し、検出精度を向上させた。
著者: Xiwen Li, Xiaoya Tang, Tolga Tasdizen
分類: cs.CV, cs.RO
原文アブストラクト
Idling vehicle detection (IVD) uses surveillance video and multichannel audio to localize and classify vehicles in the last frame as moving, idling, or engine-off in pick-up zones. IVD faces three challenges: (i) modality heterogeneity between visual cues and audio patterns; (ii) large box scale variation requiring multi-resolution detection; and (iii) training instability due to coupled detection heads. The previous end-to-end (E2E) model with simple CBAM-based bi-modal attention fails to handle these issues and often misses vehicles. We propose HAVT-IVD, a heterogeneity-aware network with a visual feature pyramid and decoupled heads. Experiments show HAVT-IVD improves mAP by 7.66 over the disjoint baseline and 9.42 over the E2E baseline.