日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
画像操作検出arXiv:2606.04545v1

Impostor: 現実的なAIGC操作位置特定のためのエージェントキュレーション型ベンチマーク

Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

シェア:XThreadsFacebookLINEはてブBluesky

生成画像編集の進歩に伴い、既存の画像操作検出・位置特定ベンチマークの限界を克服するため、100K枚の操作画像を含む新しいデータセットImpostorを構築し、閉ループエージェントフレームワークCraftAgentで多様で現実的な操作画像を自動生成した。また、局所位相モデリングと意味的フォレンジック一貫性学習を導入したPhaseAware-Netを提案し、既存手法を上回る性能を示した。

著者: Zhenliang Li, Yutao Hu, Qixiong Wang, Wenpeng Du, Hongxiang Jiang, Jiasong Wu, Xiaolong Jiang, Jungong Han

分類: cs.CV

原文アブストラクト

Recent advances in generative image editing have improved the realism and controllability of localized image manipulation, raising new challenges for image manipulation detection and localization (IMDL). However, existing IMDL benchmarks still have limitations in visual realism, manipulation diversity, and generator coverage, making it difficult to reflect recent trends in image manipulation. To address these limitations, we introduce Impostor, a high-quality AI-edited image manipulation localization dataset containing 100K manipulated images. Impostor is constructed by CraftAgent, a closed-loop agent framework that integrates scene perception, editing planning, manipulation execution, quality validation, and iterative reflection to automatically generate diverse and visually realistic manipulated images. Moreover, Impostor contains images generated by seven recent AIGC models across three manipulation types and includes multiple manipulated regions, providing a more comprehensive benchmark for AIGC-based IMDL. Furthermore, we propose PhaseAware-Net (PANet), a semantic-forensic framework that introduces local phase modeling and semantic-forensic consistency learning to better localize semantically plausible yet forensically disrupted manipulated regions. Extensive experiments show that Impostor poses significant challenges to existing large vision-language models (LVLMs) and specialized IMDL methods, while PANet achieves superior performance on Impostor and multiple public benchmarks.