ギリシャ伝統音楽のためのポーズ認識マルチモーダル自動タグ付け
Pose-Aware Multimodal Automatic Tagging for Greek Traditional Music
ギリシャ伝統音楽の自動タグ付けにおいて、音声に加えてダンサーのポーズ情報を活用するマルチモーダル手法を提案し、音声のみより約4ポイントの性能向上を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos
分類: cs.SD, cs.CV
原文アブストラクト
Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inherently multimodal, as semantic labels such as instruments, regional styles, and dance forms are encoded simultaneously across acoustic, visual, and embodied performance cues. This is especially true of culturally specific repertoires such as Greek traditional music, which remain underrepresented in MIR benchmarks. In this paper, we investigate whether the use of dancer pose provides complementary information for automatic tagging in Greek traditional music beyond audio. Using the Lyra dataset, we extend prior audio-only work by extracting aligned video features and pose-derived skeleton streams, enabling an experimental setting for multimodal auto-tagging. We further introduce an automated pipeline for extracting primary-dancer skeleton sequences from in-the-wild dance footage, combining dance-scene detection, multi-person tracking, dancer selection, pose estimation, and quality filtering. We compare unimodal, all bimodal combinations, and trimodal systems using multiple fusion strategies. Audio remains the strongest single modality (AST: macro ROC-AUC 0.821), while skeletons, though weak in isolation, enhance performance through multimodal fusion. The best trimodal system improves macro ROC-AUC by about 4 percentage points over the strongest audio baseline.
関連論文
- 大規模基盤モデルにおける音声視覚知能マルチモーダル
- 音響フィールド動画によるマルチモーダルなシーン理解マルチモーダル