日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ディープフェイク検出arXiv:2610.03380

明示的なフォレンジック特徴と時間モデリングによる動画ディープフェイクの解釈可能な検出

Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

シェア:XThreadsFacebookLINEはてブBluesky

動画から顔の時系列軌跡を抽出し、光学的・テクスチャ・幾何・圧縮の4領域にわたる68個の解釈可能な特徴をLSTMで時間モデリングすることで、高い精度と汎化性能を実現したディープフェイク検出手法。

詳しい要約

1. どんなもの?

- ビデオのDeepfake検出フレームワーク - 空間的・時間的に一貫した顔特徴をモデル化 - 解釈可能な検出を目指す - 明示的なforensic cuesを符号化 - ビデオをidentity-consistent facial trajectoriesに変換 - 固定長のtemporal windowsに分割 - 各フレームを68のstructured descriptorsで表現 - 4領域: photometric, textural, geometric, compression-based - LSTMで時間依存性を捕捉

2. 先行研究と比べてどこがすごい?

- 従来のend-to-end deep modelsは暗黙的表現に依存 - 提案手法は物理的に根拠のあるforensic cuesを明示的に符号化 - 透明な分析が可能 - マルチデータセット汎化が改善 - 解釈可能性を提供 - フレームレベルで視覚的に一貫していても時間的不整合を検出

3. 技術・手法の肝は?

- ビデオをidentity-consistent facial trajectoriesに変換 - 固定長のtemporal windowsにセグメント化 - 各フレームを68のstructured descriptorsで表現 - 4つの相補的領域: photometric, textural, geometric, compression-based - コンパクトなマルチドメイン表現 - LSTMネットワークで時間依存性と微細な不規則性を捕捉

4. どうやって有効だと検証した?

- 4つのベンチマークデータセットで評価 - FaceForensics++, Celeb-DF v2, DFDCのキュレートサブセット, DeeperForensics - F1スコア: 98.0%, 91.0%, 97.6%, 96.2% - クロスデータセット汎化も良好 - 堅牢で解釈可能なソリューションを提供

5. 議論はある?

- 要旨からは不明 - 限界や失敗ケースについての議論は記載なし - 計算コストやリアルタイム性の議論なし - 他のforensic cuesとの比較なし

6. 次に読むべき論文は?

- FaceForensics++ - Celeb-DF v2 - DeepFake Detection Challenge (DFDC) - DeeperForensics - LSTMベースの時系列モデリング - end-to-end deep models

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chahira Benhama, Mohand Saïd Allili, Assia Hamadene

分類: cs.CV

原文アブストラクト

Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.

関連論文

PR本紙発行元 EmplifAI