日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダル/誤情報検出arXiv:2608.19212v1

NepOOC-M: ネパール語-英語バイリンガルベンチマークとOOC検出のためのマルチモーダルアーキテクチャ比較分析

NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection

シェア:XThreadsFacebookLINEはてブBluesky

ネパール語を主とする初の公開OOC誤情報ベンチマークNepOOCを構築し、5つのマルチモーダルモデルとテキスト/画像のみのベースラインを比較。現段階ではテキストのみのモデルが最良で、画像のみは偶然レベルに留まることを示した。

著者: Sanjeev Khatiwada

分類: cs.CL, cs.CV

原文アブストラクト

Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC misinformation in Nepal, no public benchmark exists for Nepali. We introduce NepOOC, the first publicly available Nepali-dominant multilingual OOC benchmark, comprising 1,090 image-caption pairs (545 pristine, 545 OOC) annotated across five typologies (fabricated, miscaptioned, temporal mismatch, geographic mismatch, identity mismatch) with inter-annotator agreement kappa = 0.84. Systematic evaluation of five multimodal architectures alongside text-only and image-only baselines reveals that caption semantics appear sufficient for strong performance at the current dataset scale. A text-only mBERT model achieves 94.65+/-0.20% Macro-F1, statistically equivalent to the best multimodal system (ResNet-50+mBERT, 94.65+/-0.20%; McNemar median p = 1.000, 0/5 seeds significant at alpha = 0.05). Image-only models perform near chance (33-50%), while training-size scaling suggests that dataset expansion is a more direct path to progress than architectural sophistication or regional specialisation.