ワールドモデルベース動画生成における物理整合性の参照不要評価
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
生成動画の物理的整合性を参照データなしで評価する手法を提案し、シミュレーションと実世界のギャップを縮小する。
著者: Yun Oh, Sukmin Yun
分類: cs.AI, cs.LG, cs.RO
原文アブストラクト
We introduce reference-free measures for evaluating the physical consistency of generated videos, combining relative and absolute approaches to assess fidelity. Although tools like WorldGym or WorldEval enable robotic simulation via video generation, physical fidelity gaps often prevent these environments from accurately reproducing real-world task success rates of VLA models. Unlike existing evaluation methods, which require costly human voting (Elo) or unavailable ground-truth references (FVD), our approach utilizes DROID-SLAM and SEA-RAFT to quantify physical inconsistencies, motivated by WorldScore. Videos filtered using our relative consistency assessment show an improvement in task success rates of over 8%, effectively narrowing the simulation-to-reality gap. Furthermore, our absolute assessment enables spatio-temporal localization, providing visualization of when and where physical artifacts occur.