日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
説明可能性/音声arXiv:2606.14466v1

音声モデルにおける説明の脆さ:予測を変えずに帰属を操作する

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

シェア:XThreadsFacebookLINEはてブBluesky

音声ディープフェイク検出モデルにおいて、聞こえない摂動を用いて説明ヒートマップを操作できることを示した論文。

著者: Piotr Kitłowski, Dominik Wiącek, Mateusz Modrzejewski

分類: cs.SD, cs.AI, cs.LG

原文アブストラクト

This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. While previous work on explanation manipulation focused on images using standard $L_p$ metrics, we introduce a psychoacoustic framework that optimizes inaudible perturbations to decouple model attributions from final classifications. We evaluate this vulnerability across state-of-the-art architectures under strict prediction-preserving constraints. By evaluating the manipulation cost through domain-specific perceptual audio quality metrics alongside explanation alignment criteria, our framework demonstrates that an adversary can systematically distort automated explanation heatmaps while preserving the predicted deepfake label. Full code available at: https://github.com/cncPomper/Audio-XAI