日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2607.04112v1

DynaVieW: スキーマ誘導型ワールドモデリングによる階層的視覚ダイナミクスの理解

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

シェア:XThreadsFacebookLINEはてブBluesky

マルチモーダルLLMが動画や複数画像の時間的変化を体系的にモデル化できない問題に対し、動的スキーマに基づくワールドモデルDynaVieWを提案。状態遷移系列を学習し、視覚ダイナミクスの予測とシミュレーションを実現する。

著者: Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara Raaghavi Bhattacharya, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji, Li Mi, Syrielle Montariol, Antoine Bosselut

分類: cs.LG, cs.AI, cs.CL, cs.CV

原文アブストラクト

Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW's understanding of visual dynamics boosts its downstream performance in visual narrative creation and world simulation, showing improved consistency, controllability, and instruction-following.

関連論文