日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音声セキュリティarXiv:2606.08678v1

勾配反転と変分情報ボトルネックによる話者不変のなりすまし検出の表現学習

Speaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

シェア:XThreadsFacebookLINEはてブBluesky

話者バイアスによる汎化性能低下を解決するため、話者ラベルなしで話者不変ななりすまし検出を行う教師-学生フレームワークを提案した。

著者: Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans

分類: cs.SD, cs.LG

原文アブストラクト

Sophisticated generative speech technology can undermined the reliability of voice biometrics. While spoofing detection systems excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. In this paper, we show that such issues could be caused by speaker bias, where models learn individual voice traits rather than markers of manipulation or generation. We propose a teacher-student framework for speaker-invariant spoofing detection that disentangles identity without requiring speaker labels. We leverage a pre-trained speaker recognition teacher to guide a student model via a gradient reversal layer. To control the balance between suppressing cues related to voice identity with the preservation of those related to spoofing detection, we integrate a Variational Information Bottleneck. Evaluations across nine datasets show our model achieves a 25.7% relative reduction to the EER compared to the MHFA baseline.