DBCF: 基盤モデルの二枝相補融合による汎用ディープフェイク検出
DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection
CLIPとDINOv3という2つの基盤モデルを階層的に組み合わせ、大域的な意味情報と局所的な顔構造の手がかりを相補的に融合することで、未知の偽造手法にも汎化しやすいディープフェイク検出手法を提案した論文。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb, Shiguang Shan
分類: cs.CV
原文アブストラクト
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
関連論文
- データ多様性、周波数不変性ではない:圧縮ロバストなディープフェイク検出の制御・自己監査研究ディープフェイク検出
- 特徴ロバスト拡張と根拠に基づく説明最適化による説明可能なディープフェイク検出ディープフェイク検出
- FairForensics: 視覚言語モデルによる表情認識と人口統計解析を用いた汎化可能な公平なディープフェイク検出ディープフェイク検出
- 不確実性を考慮したマルチビュー構造学習によるディープフェイク検出ディープフェイク検出
- 説明可能なディープフェイク検出チャレンジディープフェイク検出
- 継続進化型ディープフェイク検出:動的検出システムのアーキテクチャと公開ベンチマーク評価ディープフェイク検出