セグメンテーション誘導型空間インデキシングによる汎用的かつ説明可能なディープフェイク検出
Segmentation-Guided Spatial Indexing for Generalizable and Explainable Deepfake Detection
顔のパッチトークンを意味ラベルで選別してから分類する空間インデキシング手法を導入し、口領域のみを用いた線形プローブでディープフェイクを高精度に検出する。
著者: Izaldein Al-Zyoud, Abdulmotaleb El Saddik
分類: cs.CV, eess.IV
原文アブストラクト
We introduce segmentation-guided spatial indexing for generalizable and explainable deepfake detection. The key idea reverses the standard design order: rather than pooling all facial tokens and classifying afterward, we first select semantically meaningful patch tokens, then pool only those. A frozen FaRL parser assigns each DINOv3 ViT-L/16 patch token a semantic label; non-target tokens are discarded; a linear probe classifies the retained region. This spatial indexing exploits DINOv3's patch-level spatial consistency, the same property that enables emergent segmentation, to present the probe with a purer regional subspace where manipulation-relevant evidence is less diluted by whole-face cues. Region attribution is structural: when the mouth model predicts fake, the decision used only mouth tokens, not an overlaid saliency map. On Celeb-DF v2, the mouth-indexed probe achieves AUC 0.905, outperforming LipForensics (+8.1 pp) and Xception (+16.9 pp), with no DINOv3 or FaRL fine-tuning and no target-domain data. Ablations isolate the mechanism: replacing regional selection with DINOv3's CLS token drops Celeb-DF v2 AUC by 26.4 pp; replacing DINOv3 with FaRL features drops it by 20.9 pp. Both DINOv3 representation and the spatial index are independently necessary; neither alone approaches the full system.
関連論文
- データ多様性、周波数不変性ではない:圧縮ロバストなディープフェイク検出の制御・自己監査研究ディープフェイク検出
- 特徴ロバスト拡張と根拠に基づく説明最適化による説明可能なディープフェイク検出ディープフェイク検出
- FairForensics: 視覚言語モデルによる表情認識と人口統計解析を用いた汎化可能な公平なディープフェイク検出ディープフェイク検出
- 不確実性を考慮したマルチビュー構造学習によるディープフェイク検出ディープフェイク検出
- 説明可能なディープフェイク検出チャレンジディープフェイク検出
- 継続進化型ディープフェイク検出:動的検出システムのアーキテクチャと公開ベンチマーク評価ディープフェイク検出