LoCC: 反事実フレーム整合性によるリップシンクディープフェイクの検出と位置特定
LoCC: Detection and Localization of Lip-Syncing Deepfakes via Counterfactual Frame Consistency
リップシンクディープフェイクの検出と位置特定を行う新しいフレームワークLoCCを提案。時間的近傍から生成した反事実推定との整合性を評価し、フレームレベルの不一致を捉える。
著者: Soumyya Kanti Datta, Shan Jia, Siwei Lyu
分類: cs.CV
原文アブストラクト
Lip-syncing deepfakes are among the most challenging forms of manipulated media because their artifacts are localized almost exclusively to the mouth region and evolve dynamically over time. Detecting such deepfakes requires precise temporal and spatial modeling of lip motion. In this paper, we propose LoCC, a novel detection framework that performs fine-grained detection and localization of lip-syncing deepfakes at both segment and frame levels. Unlike prior approaches that analyze videos holistically, our method evaluates whether each frame aligns with a counterfactual estimate generated from its temporal neighbors. Real videos exhibit strong and stable consistency, whereas lip-sync deepfakes introduce localized inconsistencies. Following a teacher-student learning paradigm, our model effectively captures these frame-level discrepancies and achieves superior performance over state-of-the-art methods on multiple benchmark lip-syncing deepfake datasets, including LAV-DF, AVDF1M, FakeAVCeleb, and KODF, and generalizes well across compression levels and datasets.
関連論文
- データ多様性、周波数不変性ではない:圧縮ロバストなディープフェイク検出の制御・自己監査研究ディープフェイク検出
- 特徴ロバスト拡張と根拠に基づく説明最適化による説明可能なディープフェイク検出ディープフェイク検出
- FairForensics: 視覚言語モデルによる表情認識と人口統計解析を用いた汎化可能な公平なディープフェイク検出ディープフェイク検出
- 不確実性を考慮したマルチビュー構造学習によるディープフェイク検出ディープフェイク検出
- 説明可能なディープフェイク検出チャレンジディープフェイク検出
- 継続進化型ディープフェイク検出:動的検出システムのアーキテクチャと公開ベンチマーク評価ディープフェイク検出