日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画フォレンジックarXiv:2609.30934

ManiVid: 改変動画の統合的解釈可能なフォレンジック解析

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

シェア:XThreadsFacebookLINEはてブBluesky

改変動画の検出・改変領域の特定・異常説明を統合的に行うタスクとデータセット、フレームワークを提案。

詳しい要約

1. どんなもの?

本論文は、AI生成ビデオ(AIGV)による欺瞞的な操作ビデオのリスク増大に対処するため、操作ビデオの統一的なフォレンジック分析タスクを提案する。 - タスクは、偽造検出、アーティファクトのグラウンディング、異常説明の3つを統合。 - データセットManiVid-38Kを構築。約19Kの手動検証済みリアル-フェイクビデオペアを含み、主に1080P解像度。 - 2つのパラダイムと15の強力な生成モデルで生成。 - 1KペアをManiVidBenchとしてサンプリングし、6つの操作タイプと生成モデル間でバランスを取る。 - フレームワークManiVidLensを提案。Forensic Evidence RouterとPrompt Distill Moduleを備える。

2. 先行研究と比べてどこがすごい?

先行研究と比べて以下の点が優れている。 - 既存のビデオ偽造研究は、操作ビデオに特化した高品質データセットとベンチマークが不足していた。 - MLLMは二値分類を超えるが、低レベルのフォレンジック手がかりの利用とピクセルレベルのグラウンディングが困難だった。 - 本論文は、初のデータセットManiVid-38Kを構築し、ペアのオープン語彙局所操作と真正性ラベル、偽造マスク、異常説明を組み合わせた。 - ManiVidLensは、アーティファクトグラウンディングで+21.1% mIoU、+21.3% J&F、異常説明で+131.3% ROUGE-L、+9.9% CSSの相対的改善を達成。 - 偽造検出は専用分類器と同等(0.914 Acc; 0.913 F1)。

3. 技術・手法の肝は?

技術や手法の肝は以下の通り。 - ManiVidLensは、Forensic Evidence Routerがマルチモーダル推論とビデオセグメンテーションに共有される低レベルフォレンジック証拠を提供。 - Prompt Distill Moduleがグラウンディング状態を意味的および幾何学的プロンプトに変換し、マスクデコーディングと全ビデオ伝播のための空間事前分布を蒸留。 - これにより、低レベル手がかりと高レベル推論を統合し、ピクセルレベルのグラウンディングを実現。

4. どうやって有効だと検証した?

有効性の検証は以下の通り。 - ManiVidBenchを用いて評価。1Kペアが6つの操作タイプと生成モデル間でバランスされている。 - アーティファクトグラウンディングで+21.1% mIoU、+21.3% J&Fの相対的改善。 - 異常説明で+131.3% ROUGE-L、+9.9% CSSの相対的改善。 - 偽造検出は0.914 Acc、0.913 F1で専用分類器と同等。

5. 議論はある?

議論は以下の点にある。 - 操作ビデオは完全合成ビデオと異なり、ほとんどのソースコンテンツを保持し局所領域のみを変更するため、フォレンジック分析が特に困難。 - 既存のMLLMは低レベルフォレンジック手がかりの利用とピクセルレベルのグラウンディングに苦戦。 - 本論文はデータと方法論の両面での制限に対処。 - ただし、要旨からは具体的な限界や今後の課題についての詳細は不明。

6. 次に読むべき論文は?

次に読むべき論文は以下の通り。 - 要旨で参照/比較されている研究: 既存のビデオ偽造研究、MLLMを拡張した偽造分析手法。 - 関連手法: ビデオフォレンジック、マルチモーダル大規模言語モデル(MLLM)、ビデオセグメンテーション、偽造検出。 - 具体的な論文名は要旨に記載されていないため、同分野の定番として、ビデオ偽造検出やMLLMベースのフォレンジック分析に関する論文を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li

分類: cs.CV

原文アブストラクト

Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% J&F) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).

関連論文

PR本紙発行元 EmplifAI