日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.09514

STRIKE: 物理世界モデリングのための視覚状態遷移学習

STRIKE: Learning Visual State Transitions for Physical World Modeling

シェア:XThreadsFacebookLINEはてブBluesky

現在の画像と遷移記述・経過時間から次のシーン状態を予測する遷移モデルを学習し、VLMプランナと組み合わせて物理的に整合した動画生成を実現するフレームワーク。

詳しい要約

1. どんなもの?

- 物理世界モデリングのためのフレームワーク STRIKE を提案。 - 視覚的な状態遷移学習と密な動画生成を分離する。 - 現在画像・局所遷移仕様・経過時間から次状態を予測する image-based transition model を学習。 - 推論時は pretrained vision-language planner が時間遷移仕様を予測し、遷移モデルを再帰適用して将来の視覚状態列を生成。 - 別途学習した dynamic model が状態と時間位置を条件に完全な rollout を生成。

2. 先行研究と比べてどこがすごい?

- 従来の動画生成ベースの物理世界モデリングと異なり、密な動画生成から視覚状態遷移学習を分離。 - 単に整合的な動きを生成するのではなく、相互作用がシーンをどう変えるかを予測することを目指す。 - 対応する video-backbone baselines に対し、物理的一貫性と manipulation-video fidelity のベンチマーク指標で改善。 - 学習された視覚状態遷移が物理世界モデリングの有効な中間表現であることを支持。

3. 技術・手法の肝は?

- 訓練動画から観測状態を抽出し、遷移記述と時間オフセットを対にした event-aligned supervision を構築。 - image-based transition model が現在画像・局所遷移仕様・経過時間から次シーン構成を予測。 - 推論では pretrained vision-language planner が時間遷移仕様を予測。 - 学習済み遷移モデルを再帰適用し、将来の視覚状態列を生成。 - 別途訓練した dynamic model が状態と時間位置を条件に完全な rollout を生成。

4. どうやって有効だと検証した?

- Physics-IQ Verified、PhyGenBench、Pisa-Experiments、RoboTwin2.0 で実験。 - 対応する video-backbone baselines と比較。 - 物理的一貫性と manipulation-video fidelity のベンチマーク指標で STRIKE の改善を確認。

5. 議論はある?

- 学習された視覚状態遷移が物理世界モデリングの有効な中間表現であると結論。 - 限界や失敗事例、計算コスト、一般化性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている video-backbone baselines や、Physics-IQ Verified、PhyGenBench、Pisa-Experiments、RoboTwin2.0 に関連する研究。 - 具体的な論文名は要旨からは不明。 - 同分野の定番として video generation、vision-language planning、world model に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenbin Teng, Tianshuo Xu, Depu Meng, Yuelei Li, Quentin Herau, Yihan Hu, Yajie Zhao, Wei Zhan

分類: cs.CV

原文アブストラクト

Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

関連論文

PR本紙発行元 EmplifAI