日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画世界モデルarXiv:2609.18430

StrucPhysVideo: 構造化キャプションとロボット行動から物理ダイナミクスを学習する動画世界モデル

StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions

シェア:XThreadsFacebookLINEはてブBluesky

物体の動きや相互作用を記述した構造化キャプションとロボットの行動を条件として、物理的に妥当な動画を生成・予測する動画世界モデルを提案し、Physics-IQで最高性能を達成した。

詳しい要約

1. どんなもの?

- 物理ダイナミクスを扱う video world model 群 - 物体の運動・相互作用・状態変化を予測 - 対象は embodied AI 向け - 構成要素は2つ - 物理重視のデータキュレーション - 言語・行動条件付きの scene evolution 予測 - モデル系列 - StrucPhysVideo-TI2V: text-image-to-video - StrucPhysVideo-IA2V: image-action-to-video - データは構造化キャプション付き - objects, materials, 時間局所的な interactions - contact, deformation, state transitions を明示

2. 先行研究と比べてどこがすごい?

- Physics-IQ Verified で SOTA - 45.5% を記録 - Cosmos3-Super-Image2Video を 2.8 ポイント上回る - 物理重視の監督信号を明示的に設計 - caption ablations で有効性を確認 - 単なる image/language 条件付き予測を超える - robot end-effector commands による action-driven interaction へ拡張 - 先行研究との詳細な比較は要旨からは不明

3. 技術・手法の肝は?

- データパイプライン - motion-aware video segmentation - quality/content filtering - physical relevance verification - objects, materials, temporally localized interactions の構造化注釈 - カメラ運動と物体挙動を分離 - contact, deformation, state transitions を記述 - StrucPhysVideo-TI2V - sparse Mixture-of-Experts (MoE) - curriculum 学習で物理ダイナミクスを段階的に重視 - 一般ドメイン動画データも保持 - StrucPhysVideo-IA2V - action conditioning - causal autoregressive generation - few-step distillation で4 denoising steps の逐次 rollout

4. どうやって有効だと検証した?

- Physics-IQ Verified で評価 - StrucPhysVideo-TI2V が 45.5% を達成 - Cosmos3-Super-Image2Video を 2.8 ポイント上回る - caption ablations を複数 backbone で実施 - 物理重視 supervision の有効性を確認 - StrucPhysVideo-IA2V の検証 - robot end-effector commands からの visual outcome 予測 - 4 denoising steps での incremental robot rollouts - 詳細な実験設定は要旨からは不明

5. 議論はある?

- 物理ダイナミクスモデリングの進展を主張 - image/language 条件付きから action-driven interaction へ - データキュレーションと構造化注釈の重要性を示唆 - 限界や失敗事例の議論は要旨からは不明 - 計算コストやスケーラビリティの議論は要旨からは不明

6. 次に読むべき論文は?

- Cosmos3-Super-Image2Video - 比較対象として明示 - Physics-IQ Verified - 評価ベンチマークとして参照 - 関連手法 - video world models - Mixture-of-Experts (MoE) - text-image-to-video - image-action-to-video - causal autoregressive generation - few-step distillation

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Awomo-WM Team, :, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu

分類: cs.CV

原文アブストラクト

Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucPhysVideo, a family of video world models that bridges physics-focused data curation with language- and action-conditioned prediction of scene evolution. Our data pipeline combines motion-aware video segmentation, quality and content filtering, and physical relevance verification with structured annotations of objects, materials, and temporally localized interactions. By disentangling camera motion from object behavior and explicitly describing contact, deformation, and state transitions, the pipeline provides supervision grounded in observable physical events. Building on these data, we introduce StrucPhysVideo-TI2V, a sparse Mixture-of-Experts (MoE) text-image-to-video model trained with a curriculum that progressively emphasizes physical dynamics while retaining general-domain video data. StrucPhysVideo-TI2V achieves state-of-the-art performance on Physics-IQ Verified, scoring 45.5% and outperforming Cosmos3-Super-Image2Video by 2.8 percentage points. Caption ablations across backbones further demonstrate the effectiveness of physics-focused supervision. We further extend StrucPhysVideo-TI2V to StrucPhysVideo-IA2V, an interactive image-action-to-video world model that predicts visual outcomes from robot end-effector commands. Action conditioning, causal autoregressive generation, and few-step distillation enable incremental robot rollouts with only four denoising steps. Together, StrucPhysVideo advances physical dynamics modeling from image- and language-conditioned video prediction toward action-driven interaction.

関連論文

PR本紙発行元 EmplifAI