日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38057

EVO-WAM: 映像と行動の検証による世界行動モデルの進化

EVO-WAM: Evolving World Action Models through Video-Action Verification

シェア:XThreadsFacebookLINEはてブBluesky

世界行動モデルが生成した映像と行動の軌跡を自己検証しながら反復学習し、追加の実演データなしで未知タスクへの適応を可能にしたフレームワーク。

詳しい要約

1. どんなもの?

本論文は、追加のexpert demonstrationsを収集せずに新しいタスクへrobot policiesを適応させるためのフレームワークEVO-WAMを提案する。World action models (WAMs) が生成するvideo-action trajectoriesを自己学習に用いる。外部環境で候補行動を実行せず、生成されたrolloutsから学習する。

2. 先行研究と比べてどこがすごい?

従来のWAM適応は、生成videoがタスク完了を描けない、または見た目が成功でもactionが不整合で実行失敗する問題があった。EVO-WAMは外部実行フィードバックなしで自己生成軌跡から学習し、vision-language modelとinverse dynamics modelで信頼できる経験を選別する点が新しい。

3. 技術・手法の肝は?

- WAM trainingにstate predictionとanchored multi-frame contextを追加し、外部実行フィードバックなしで完全なautoregressive rolloutsを可能にする。 - vision-language modelでtask-completing prefixesを選択し、inverse dynamics modelでvideo-action consistencyを検証して信頼できるtraining experienceを同定する。 - 検証済みprefixesでWAMを反復訓練し、更新モデルで新rolloutsを生成する。

4. どうやって有効だと検証した?

- 7つの未見RoboTwin 2.0タスクで、Cosmos3の平均成功率を26.9%から68.0%へ、DreamZeroを28.5%から46.4%へ向上(約2.5倍と1.6倍)。 - 実世界の3つの未見long-horizon composite tasksで、Cosmos3の平均成功率を20.0%から76.7%へ改善(56.7ポイント増)。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究として、World action models (WAMs)、Cosmos3、DreamZero、RoboTwin 2.0が挙げられる。関連手法としてvision-language model、inverse dynamics modelも参照されている。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian

分類: cs.CV, cs.RO

原文アブストラクト

Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.

関連論文

PR本紙発行元 EmplifAI