日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
視覚編集arXiv:2606.00188

PaintBench: 精密な視覚編集の決定的評価

PaintBench: Deterministic Evaluation of Precise Visual Editing

シェア:XThreadsFacebookLINEはてブBluesky

PaintBenchは、幾何変換や構造操作など20種類の精密な視覚編集操作を評価するベンチマークで、手続き生成とピクセル単位の評価によりバイアスを排除し、既存モデルの性能が低いことを示した。

著者: Kai Xu, Ellis Brown, Shrikar Madhu, Rob Fergus, He He, Saining Xie

分類: cs.GR, cs.CV, cs.LG

原文アブストラクト

While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To probe this challenge, we introduce PaintBench, a dynamically scalable benchmark targeting 20 fundamental precise visual editing operations across four categories: geometric transformation, structural manipulation, color change, and symbolic reasoning. Procedural generation with configurable complexity enables an effectively infinite, contamination-resistant evaluation suite, and deterministic pixel-level evaluation eliminates reliance on bias-prone judge models. Across 11 image editing models, we find overall low performance, with the current highest-performing industry leader scoring only 17.1% (mIoU). Task decomposition reveals especially challenging operation types (geometric transformation, most structural manipulation, formula-based color change) and model-specific specializations. Fine-grained benchmark diagnostics further show performance degradations induced by scene variations in object count, background complexity, color scheme, and edit-region size. To test generalization of PaintBench scores to applied task performance, we create a procedural, deterministic evaluation for data visualization editing (TinyGrafixBench) and find strong linear correlation with PaintBench scores ($R^2 = 0.91$, $p < 0.001$). Altogether, PaintBench provides a rigorous foundation for measuring and driving progress in precise multimodal visual editing.