日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10270

動画予測ポリシー2:より良く予測し、より良く行動する

Video Prediction Policy 2: Predict Better, Act Better

シェア:XThreadsFacebookLINEはてブBluesky

操作動画で基盤動画モデルを追加事前学習し、単段階の視覚プランナに蒸留、MoTで行動モジュールを統合することで、オープンエンドな操作タスクにおけるゼロショット汎化性能を大幅に向上させた世界行動モデル。

詳しい要約

1. どんなもの?

- 汎用ロボットポリシーとしての World action models (WAMs) の一種 - 動画予測の事前知識を行動学習に転移することを目指す - 既存 WAMs は open-ended 環境で誤った動き予測をし、誤行動を生む問題を指摘 - 原因を (1) base video models が manipulation 向けに最適化されていない、(2) 行動成分の安易な組み込みが汎化を損なう、と分析 - 提案手法 Video Prediction Policy 2 (VPP2) は video prediction と action generation の両方で強い zero-shot 汎化を実現

2. 先行研究と比べてどこがすごい?

- 既存 WAMs は open-ended 環境で不正確な動き予測をし、誤行動につながる - VPP2-14B は Cosmos3-64B を open-ended タスクの video prediction instruction-following 成功率で 11.0% ポイント上回る - 実世界 zero-shot ALOHA manipulation タスクで最強 baseline を 18.5% ポイント上回る - LIBERO-Pro, LIBERO-OOD, RoboDojo ベンチマークで評価手法中最高の成功率を達成

3. 技術・手法の肝は?

- 大規模で多様な manipulation 動画データセットを収集し、base video foundation model を継続事前学習 - 動画クリップに詳細キャプションを付与し、event-level 動画事前学習を実施し open-ended manipulation タスクへの汎化を促進 - 動画モデルを post-train および distill し、固定予測ホライズンの single-step visual planner に変換 - mixture-of-transformers (MoT) アーキテクチャで action module を導入し、暗黙的 inverse dynamics model を学習

4. どうやって有効だと検証した?

- open-ended タスクでの video prediction instruction-following 成功率を評価 - 実世界 zero-shot ALOHA manipulation タスクで成功率を評価 - LIBERO-Pro, LIBERO-OOD, RoboDojo ベンチマークでベンチマーク固有の post-training 後に成功率を評価 - 結果として VPP2-14B が Cosmos3-64B を 11.0% ポイント上回り、最強 baseline を 18.5% ポイント上回り、各ベンチマークで最高成功率を達成

5. 議論はある?

- 既存 WAMs の誤った動き予測の原因として base video models の manipulation 非最適化と行動成分組み込みによる汎化低下を指摘 - これらの課題に対処する設計を提案 - 具体的な限界や失敗事例、計算コスト、倫理的影響などの議論は要旨からは不明

6. 次に読むべき論文は?

- Cosmos3-64B (比較対象の base video model) - ALOHA (実世界 manipulation タスク) - LIBERO-Pro, LIBERO-OOD, RoboDojo (ベンチマーク) - World action models (WAMs) 関連の先行研究 - mixture-of-transformers (MoT) アーキテクチャ - inverse dynamics model 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yanjiang Guo, Haodong Yan, Zhide Zhong, Zhongru Zhang, Qingyuan Yang, Qingzhou Lu, Xiaoyu Chen, Yen-Jen Wang, Shuying Deng, Chenghan Yang, Puzhen Yuan, Chenxin Liu, Tun Ban, Xiang Zhu, Yichen Liu, Kun Feng, Haoang Li, Jianyu Chen

分類: cs.CV, cs.RO

原文アブストラクト

World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textit{event-level} video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0\% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5\% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.

関連論文

PR本紙発行元 EmplifAI