日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ディープフェイク検出arXiv:2609.34720

DBCF: 基盤モデルの二枝相補融合による汎用ディープフェイク検出

DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

シェア:XThreadsFacebookLINEはてブBluesky

CLIPとDINOv3という2つの基盤モデルを階層的に組み合わせ、大域的な意味情報と局所的な顔構造の手がかりを相補的に融合することで、未知の偽造手法にも汎化しやすいディープフェイク検出手法を提案した論文。

詳しい要約

1. どんなもの?

- 顔画像のDeepfake検出を対象とした新しいフレームワークDBCFを提案する研究。 - 複数のFoundation Modelを階層的に融合する点が特徴。 - CLIPベースのGlobal Context Branch (GCB)とDINOv3ベースのFine-grained Cue Branch (FCB)を統合。 - 大規模Foundation Modelの豊富な表現と汎化性能を活用し、未知の操作や異なるドメインへの汎化を目指す。

2. 先行研究と比べてどこがすごい?

- 従来の小規模な検出モデルは偽造手がかりの捕捉能力が限られ、ドメインや未知の操作への汎化が困難。 - 単一のFoundation Modelでは不十分:CLIPは大域的意味手がかりに優れるが局所的な顔特徴の捕捉が弱い。 - DINOは局所構造特徴の捕捉に優れるが大域的意味文脈が弱い。 - これらを相補的に融合することで、より包括的な偽造表現を学習し、クロス操作性能を向上。

3. 技術・手法の肝は?

- 階層的マルチグラニュラーフレームワークを提案。 - GCB (CLIPベース) が全体的な意味手がかりを捕捉。 - FCB (DINOv3ベース) が局所的な構造的不規則性を捕捉。 - 特徴融合モジュールを設計し、凍結したFoundation Backboneをパラメータ効率的に適応。 - 2モデルから相補的特徴を適応的に抽出・統合。

4. どうやって有効だと検証した?

- 複数のベンチマークで広範な実験を実施。 - 特にクロスデータセットおよびクロス操作設定において有効性を実証。 - 提案設計の利点を示す結果が得られた。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、倫理的影響などについての議論は要旨では触れられていない。

6. 次に読むべき論文は?

- CLIP (Radford et al.) - DINO / DINOv3 - その他のDeepfake検出におけるFoundation Model活用研究(要旨では具体的な参照論文は明示されていない)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb, Shiguang Shan

分類: cs.CV

原文アブストラクト

As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.

関連論文

PR本紙発行元 EmplifAI