日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成評価arXiv:2606.22363

ワールドモデルベース動画生成における物理整合性の参照不要評価

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

シェア:XThreadsFacebookLINEはてブBluesky

生成動画の物理的整合性を参照データなしで評価する手法を提案し、シミュレーションと実世界のギャップを縮小する。

著者: Yun Oh, Sukmin Yun

分類: cs.AI, cs.LG, cs.RO

原文アブストラクト

We introduce reference-free measures for evaluating the physical consistency of generated videos, combining relative and absolute approaches to assess fidelity. Although tools like WorldGym or WorldEval enable robotic simulation via video generation, physical fidelity gaps often prevent these environments from accurately reproducing real-world task success rates of VLA models. Unlike existing evaluation methods, which require costly human voting (Elo) or unavailable ground-truth references (FVD), our approach utilizes DROID-SLAM and SEA-RAFT to quantify physical inconsistencies, motivated by WorldScore. Videos filtered using our relative consistency assessment show an improvement in task success rates of over 8%, effectively narrowing the simulation-to-reality gap. Furthermore, our absolute assessment enables spatio-temporal localization, providing visualization of when and where physical artifacts occur.