日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダル検証arXiv:2609.31766

UNMATCH: 画像と主張の対応検証のための選択的不均衡トークン・パッチマッチング

UNMATCH: Selective Unbalanced Token-Patch Matching for Forensic Image-Claim Verification

シェア:XThreadsFacebookLINEはてブBluesky

画像とテキスト主張の局所的な不一致を方向性マルチスケール被覆度で捉え、軽量な分類器で偽情報ペアを検出する手法を提案。

著者: Xinjin Li, Lian Lian, Yuanzhe Yang, Yudi Xia, Calvin Chang Liu, Yeyun Xu, Yu Ma, Jinghan Cao, Yuruo Gong

分類: cs.CV

原文アブストラクト

Contextual image misuse pairs an image with a misleading claim. We study image-claim correspondence in fact-checked pairs containing out-of-context reuse, visual manipulation, or both. Existing pair-based detectors often compress the two modalities into a global compatibility score or learn a highly flexible interaction module, which can obscure a decisive local mismatch. We introduce directional multiscale coverage, a compact representation that summarizes local image-claim affinity in both directions and at three spatial scales. At each scale, each direction is summarized by its mean, lower quartile, and two thresholded support ratios; the signed difference between directional means completes a nine-dimensional scale descriptor. Concatenating the three scales yields a compact local representation for a lightweight global-local classifier. Under leakage-aware three-fold, three-seed evaluation on the Snopes subset of the Fauxtography benchmark, UNMATCH achieves 69.82 Macro-F1 and 71.05 balanced accuracy, exceeding the MCOT adaptation by 2.60 and 2.16 points. A matched-reassigned intervention shows that breaking the observed pairing lowers coverage and increases both discrepancy and false-pair probability.

PR本紙発行元 EmplifAI