日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ビデオ生成/世界モデルarXiv:2603.28489

世界モデルとしてのビデオ生成モデル:効率的なパラダイム、アーキテクチャ、アルゴリズム

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

シェア:XThreadsFacebookLINEはてブBluesky

ビデオ生成モデルを世界シミュレータとして実用化するための効率性に焦点を当て、効率的なモデリングパラダイム、ネットワークアーキテクチャ、推論アルゴリズムの3次元で体系的なレビューを提供する。

著者: Muyang He, Hanzhong Guo, Junxiong Lin, Yizhou Yu

分類: eess.IV, cs.CV

原文アブストラクト

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical capacity for world simulation and the heavy computational costs of spatiotemporal modeling. To address this, we comprehensively and systematically review video generation frameworks and techniques that consider efficiency as a crucial requirement for practical world modeling. We introduce a novel taxonomy in three dimensions: efficient modeling paradigms, efficient network architectures, and efficient inference algorithms. We further show that bridging this efficiency gap directly empowers interactive applications such as autonomous driving, embodied AI, and game simulation. Finally, we identify emerging research frontiers in efficient video-based world modeling, arguing that efficiency is a fundamental prerequisite for evolving video generators into general-purpose, real-time, and robust world simulators. A curated GitHub repository of the reviewed literature can be found at https://github.com/Isaachhh/Efficient-VWM-Survey.