日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12417

WOVEN: 視覚的遷移推論をマルチモーダルLLMに織り込む

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

シェア:XThreadsFacebookLINEはてブBluesky

視覚的遷移推論を共通の学習プリミティブとして扱い、そのための学習データとベンチマークWOVENを構築。38の最先端MLLMが人間を大きく下回ることを示し、WOVENで訓練すると広範なタスクへ転移することを実証した。

詳しい要約

1. どんなもの?

視覚的遷移推論(visual transition reasoning)をMLLMsの共通訓練プリミティブとして扱う研究。WOVENという訓練ソース兼ベンチマークを導入し、video-pretrained generative modelsによる多様で現実的なrolloutから、36,076例を20 scene types・5 action types・8 reasoning typesで整理。38のfrontier MLLMsを評価し、人間を大きく下回る体系的欠陥を確認。複数スケールでWOVEN訓練し、広く転移する共有能力の獲得を示す。

2. 先行研究と比べてどこがすごい?

既存ベンチマークは空間・身体・物理・時間推論の欠陥を個別に記録するが、scene・action・reasoning operationをまたぐ統制比較を支援しない。WOVENは遷移監督をscene・action・reasoning typeで組織化し、制御された比較を可能にする。さらに訓練データとして、約2,000項目の部分集合が26外部ベンチマーク中22を最大27.3ポイント改善し、タスク自身の訓練データの30-50%を置換可能と示す。

3. 技術・手法の肝は?

video-pretrained generative modelsによる多様で現実的なrolloutを用い、遷移監督をscene・action・reasoning typeで整理したWOVENを構築。複数スケールのMLLMsをWOVENで訓練。統制比較からvisual world modelingの訓練レシピを導出: 監督はそれが教えるreasoning operationで選択し、action・scene・domainでは選ばない。robustnessにはvisual stateのより大きな変化を好む。

4. どうやって有効だと検証した?

38のfrontier MLLMs(例: GPT-5.4, Qwen3-VL-235B-A22B)を評価し、最強モデルでも人間を大きく下回り、失敗がモデルファミリーをまたぎスケールでも持続することを確認。WOVENで複数スケールのMLLMsを訓練し、約2,000項目の部分集合が26外部ベンチマーク中22を最大27.3ポイント改善、タスク自身の訓練データの30-50%を置換可能と検証。訓練レシピはheld-out benchmarksで前向きに検証。

5. 議論はある?

視覚的遷移推論がMLLMsの空間・身体・物理・時間推論の失敗に共通する欠陥であるという仮説を提示。この能力が異なる監督源から学習され、異なるタスク間で再利用可能な共有訓練プリミティブとなり得るかを検証。訓練レシピの一般性や限界、他分野への適用可能性については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として、video-pretrained generative models、MLLMs(例: GPT-5.4, Qwen3-VL-235B-A22B)、visual world modeling、visual transition reasoningのベンチマークが挙げられる。同分野の定番としてembodied AIやspatial reasoningのベンチマークも次に読む候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li

分類: cs.CV, cs.CL, cs.LG

原文アブストラクト

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.

関連論文

PR本紙発行元 EmplifAI