日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音声ディープフェイク検出arXiv:2607.21820v1

音声ディープフェイク検出器における話者同一性感度の調査

Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

シェア:XThreadsFacebookLINEはてブBluesky

音声ディープフェイク検出器が話者同一性に依存する問題を指摘し、その感度を定量化する指標ISSを提案した。

著者: Daniyal Kabir Dar, Arun Ross

分類: cs.SD, cs.AI, cs.CR, cs.LG

原文アブストラクト

Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing factor is speaker-identity reliance: standard training corpora correlate speaker identity with the genuine/synthetic label, allowing detectors to partially rely on speaker-related cues rather than synthesis artifacts alone. We propose the Identity Sensitivity Score (ISS), a per-utterance diagnostic that quantifies how much a detector's output changes across different speaker identity contexts. ISS requires no ground-truth labels at inference time and can be computed from the detector score and a pool of reference speaker examples. Across two detectors and two datasets, incorrectly classified utterances have ISS scores 29 to 52 times higher than correctly classified utterances, and ISS alone predicts misclassification with area-under-curve (AUC) up to 0.954. To test whether ISS actually captures identity-sensitive behavior rather than serving only as a proxy for prediction confidence, we apply voice conversion to 500 utterances and measure the resulting detector-score shift. Utterances flagged as identity-sensitive by ISS respond 19 to 30 times more strongly to this manipulation than utterances flagged as stable. These results position ISS as a practical inference-time diagnostic for speaker-dependent failure analysis in audio deepfake detection.