日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダルarXiv:2609.27094

ギリシャ伝統音楽のためのポーズ認識マルチモーダル自動タグ付け

Pose-Aware Multimodal Automatic Tagging for Greek Traditional Music

シェア:XThreadsFacebookLINEはてブBluesky

ギリシャ伝統音楽の自動タグ付けにおいて、音声に加えてダンサーのポーズ情報を活用するマルチモーダル手法を提案し、音声のみより約4ポイントの性能向上を達成した。

詳しい要約

1. どんなもの?

- Greek traditional music を対象に、audio に加えて dancer pose を統合する multimodal automatic tagging を検討した研究。 - Lyra dataset を拡張し、aligned video features と pose-derived skeleton streams を抽出。 - in-the-wild の dance footage から primary-dancer skeleton sequences を自動抽出する pipeline も提案。 - unimodal、全 bimodal 組み合わせ、trimodal を複数の fusion strategies で比較。

2. 先行研究と比べてどこがすごい?

- 従来の tagging は audio のみが主流で、live performance の visual/embodied cues を活用していない。 - Greek traditional music のような文化的に特定の repertoire は MIR benchmarks で過小代表。 - 先行の audio-only 研究を video features と pose-derived skeleton streams で拡張。 - skeleton は単独では弱いが、multimodal fusion で性能を向上させる点を示した。

3. 技術・手法の肝は?

- Lyra dataset を基に aligned video features と pose-derived skeleton streams を抽出。 - in-the-wild dance footage から primary-dancer skeleton sequences を自動抽出する pipeline を導入。 - pipeline は dance-scene detection、multi-person tracking、dancer selection、pose estimation、quality filtering を組み合わせる。 - unimodal、全 bimodal 組み合わせ、trimodal を複数の fusion strategies で比較。

4. どうやって有効だと検証した?

- Lyra dataset 上で unimodal、bimodal、trimodal システムを比較評価。 - Audio が最強の単一 modality (AST: macro ROC-AUC 0.821)。 - Skeleton は単独では弱いが、multimodal fusion で性能を向上。 - 最良の trimodal system は最強 audio baseline より macro ROC-AUC を約 4 percentage points 改善。

5. 議論はある?

- Audio が依然として最強の単一 modality である一方、skeleton は fusion を通じて補完的情報を提供。 - 文化的に特定の repertoire における multimodal tagging の可能性を示す。 - 限界や課題、今後の方向性についての詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 先行の audio-only work、Lyra dataset を用いた研究。 - 関連手法: AST (Audio Spectrogram Transformer)、multimodal fusion strategies、pose estimation、multi-person tracking。 - 同分野の定番: Music Information Retrieval (MIR) における automatic tagging、multimodal music tagging。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos

分類: cs.SD, cs.CV

原文アブストラクト

Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inherently multimodal, as semantic labels such as instruments, regional styles, and dance forms are encoded simultaneously across acoustic, visual, and embodied performance cues. This is especially true of culturally specific repertoires such as Greek traditional music, which remain underrepresented in MIR benchmarks. In this paper, we investigate whether the use of dancer pose provides complementary information for automatic tagging in Greek traditional music beyond audio. Using the Lyra dataset, we extend prior audio-only work by extracting aligned video features and pose-derived skeleton streams, enabling an experimental setting for multimodal auto-tagging. We further introduce an automated pipeline for extracting primary-dancer skeleton sequences from in-the-wild dance footage, combining dance-scene detection, multi-person tracking, dancer selection, pose estimation, and quality filtering. We compare unimodal, all bimodal combinations, and trimodal systems using multiple fusion strategies. Audio remains the strongest single modality (AST: macro ROC-AUC 0.821), while skeletons, though weak in isolation, enhance performance through multimodal fusion. The best trimodal system improves macro ROC-AUC by about 4 percentage points over the strongest audio baseline.

関連論文

PR本紙発行元 EmplifAI