日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ディープフェイク検出arXiv:2610.09952

MSUチームによる説明可能なディープフェイク検出チャレンジ2026への挑戦:根拠あるアーティファクト証拠に基づく検出

MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

シェア:XThreadsFacebookLINEはてブBluesky

ディープフェイク画像の真偽判定と、目に見える証拠に基づく説明文を同時に生成する手法を提案し、複数のDINOv3と操作位置特定特徴を組み合わせた検出器と、アーティファクト証拠マップによる説明生成を実現した。

詳しい要約

1. どんなもの?

- XPlainVerse dataset [1] を用いた Explainable Deepfake Detection Challenge [2] への MSU チームの解決策 - 画像の real/fake 判定と、視覚的な forensic cue に基づく complex/simple な説明生成を同時に行う - modular な detection-and-explanation 設計を採用 - 検出器は複数の DINOv3 と Mesorch の manipulation-localization 特徴を組み合わせた multi-backbone - 説明生成は class-conditional Qwen3-VL と GRPO 最適化された text simplification model を使用

2. 先行研究と比べてどこがすごい?

- 従来の deepfake 検出は精度重視で、決定の視覚的根拠を提示しないものが多い - 本手法は検出と説明を統合し、Artifact Evidence Map により検出器に説明根拠を注入 - paired images や pixel-level manipulation masks を必要とせず、weak patch-level supervision で学習可能 - 複数 backbone と DCT-aware cue、multi-scale forensic 情報を統合する点が先行研究と異なる - 要旨からは具体的な先行研究との定量比較は不明

3. 技術・手法の肝は?

- real/fake 判定: 複数 DINOv3 と Mesorch manipulation-localization 特徴を統合した multi-backbone detector - 説明根拠の注入: Grounding-DINO ベースの pseudo-mask generation pipeline で training explanations の local artifact descriptions を weak patch-level supervision に変換し Artifact Evidence Map を生成 - 特徴空間分離: paired images や pixel-level masks 不要の local patch-level contrastive objective で artifact と authenticity の証拠を分離 - 言語出力: class-conditional Qwen3-VL で complex explanations を生成し、GRPO 最適化 text simplification model で simple…

4. どうやって有効だと検証した?

- 提案手法を XPlainVerse の challenge subset で訓練・評価 - full test split で detection accuracy 0.9349、explanation score 0.5571、final challenge score 0.7456 を達成 - 要旨からは ablation study や他手法との比較の詳細は不明

5. 議論はある?

- 要旨からは議論や限界、失敗事例についての記述は不明 - 検出精度は高いが explanation score は 0.5571 と相対的に低く、説明品質に課題が残る可能性 - 要旨からは具体的な議論は不明

6. 次に読むべき論文は?

- XPlainVerse dataset [1] - Explainable Deepfake Detection Challenge [2] - DINOv3 - Mesorch - Grounding-DINO - Qwen3-VL - GRPO - 同分野の定番として Grad-CAM や LIME などの説明手法、および deepfake detection の一般的な backbone (Xception, EfficientNet など)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Artem Filippov, Aleksandr Gushchin, Kirill Koltsov, Dmitriy Vatolin, Anastasia Antsiferova

分類: cs.CV

原文アブストラクト

Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.

関連論文

PR本紙発行元 EmplifAI