日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ディープフェイク検出arXiv:2606.15117v1

音声・映像ディープフェイク検出のための教師-学生構造によるドメイン適応アンサンブルモデル

Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection

シェア:XThreadsFacebookLINEはてブBluesky

音声と映像の両方を扱うディープフェイク検出モデルを、教師-学生フレームワークでドメイン適応させ、未知のデータセットでの性能を向上させる手法を提案した。

著者: Elham Abolhasani, Maryam Ramezani, Hamid R. Rabiee

分類: cs.MM, cs.AI, cs.CV, cs.LG, cs.SD

原文アブストラクト

The rapid advancement of generative AI models is leading to more realistic deepfake media, encompassing the manipulation of audio, video, or both. This raises severe privacy and societal concerns. Numerous studies in this area have yielded promising intra-domain results; however, these models frequently exhibit decreased efficacy when faced with data from dissimilar domains. Consequently, recent deepfake detection approaches focus on enhancing the generalization ability through multiple techniques that incorporate all input modalities, including audio, images, and their interactions. In this regard, we propose the EAV-DFD method, a generalized deep ensemble audio-visual model (EAV-DFD) combined with a domain adaptation mechanism utilizing a teacher-student framework to enhance the model's ability to perform and generalize effectively across unseen domains. To evaluate the model's performance, we used the FakeAVCeleb dataset as the primary domain and the DFDC, Deepfake_TIMIT, and PolyGlotFake datasets as an unseen domain. Our experimental results demonstrate that the proposed framework is efficient in domain adaptation, improving AUC performance of the model by 4.09%, 17.94%, and 0.5% on three unseen datasets, using only a small portion of them to train the student model. This leads to a novel deepfake detection model capable of adapting to new domains and interpreting which modality has been manipulated, highlighting the potential of our approach for real-world applications.

関連論文