日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
文書鑑定/VLMarXiv:2609.38391

文書画像の改ざん検出・局所化・意味的異常検知・鑑定レポート生成を統合したCoTパイプライン

Team MSU GenText-Forensics Challenge 2026 Technical Report

シェア:XThreadsFacebookLINEはてブBluesky

文書画像の改ざん検出器と2つのLoRA適応VLMを組み合わせ、改ざん領域の特定・攻撃種別判定・意味的異常の発見・鑑定レポート生成を一貫して行う手法を提案し、ACM MM 2026 GenText-Forensicsチャレンジで3位を獲得した。

著者: Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova

分類: cs.CV

原文アブストラクト

Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR & LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.

PR本紙発行元 EmplifAI