日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02508

漸進的視覚計画による世界行動モデリング

World Action Modeling with Progressive Visual Planning

シェア:XThreadsFacebookLINEはてブBluesky

将来の視覚と行動を同時予測する世界行動モデルにおいて、疎な視覚サブゴール列を段階的に予測することで長期的タスクを効率的に計画・実行する手法ProWAMを提案。

詳しい要約

1. どんなもの?

- 提案手法: ProWAM - ロボット制御のためのWorld Action Model (WAM) - 初期観測と指示から将来の視覚ダイナミクスと行動を同時予測 - 長期的予測を効率化するため、スパースな視覚サブゴール列を逐次予測 - 行動生成と視覚計画を統合 - 目的: 長期的タスクにおけるロボット制御の性能向上 - 特にout-of-distribution (OOD) ロバスト性の改善 - 特徴: 進行的な視覚計画 (progressive visual planning) - サブゴールが行動生成のアンカーとして機能 - 大規模action-free動画から学習可能

2. 先行研究と比べてどこがすごい?

- 既存WAMの問題点 - 長期的予測が困難: 密なビデオロールアウト生成が非効率 - 一部の手法は単一未来フレームのみ予測し、目標への進展を無視 - ProWAMの優位性 - スパースな視覚サブゴール列を予測し、明示的な視覚ガイダンスを提供 - ビデオバックボーンが複雑な視覚計画をオフロードし、行動ポリシーを軽量化 - 単一のビデオバックボーン順伝播でサブゴール特徴をキャッシュし、反復的な全ビデオ生成を排除 - 再計画時は軽量な行動デノイジングのみで効率的 - シミュレーションベンチマークで新たなstate-of-the-artを達成 - LIBERO-Plus: 85.8% - randomized RoboTwin: 75.7% - 最強ベースラインに対し最大+35.9%の相対改善 - RoboCasa365: 48.1%成功、Composite-Unseenで18.2%、総合4位 - ゼロショット実世界実験: 70.0%成功、最強ベースラインを+15.0上回る (55.0%→70.0%、相対+27.3%)

3. 技術・手法の肝は?

- 核心: 進行的World Action Model (ProWAM) - 行動と順序付けられたスパース視覚サブゴール列を同時予測 - サブゴールが行動生成のアンカーとして機能 - 学習 - サブゴール予測は大規模action-free動画から学習可能 - ビデオバックボーンが視覚計画を担当し、行動ポリシーから複雑さを分離 - 推論効率 - 単一のビデオバックボーン順伝播でスパースサブゴール特徴をキャッシュ - 反復的な全ビデオ生成を不要に - 再計画時は軽量な行動デノイジングのみ - 設計のスケーラビリティ - サブゴール予測の学習が大規模動画に自然にスケール

4. どうやって有効だと検証した?

- シミュレーションベンチマーク - LIBERO-Plus: 85.8% (state-of-the-art) - randomized RoboTwin: 75.7% (state-of-the-art) - 最強ベースラインに対する相対改善: 最大+35.9% - RoboCasa365: 48.1%成功、Composite-Unseenで18.2%、総合4位 - 実世界実験 - ゼロショット設定 - 新規シーンで70.0%成功 - 最強ベースラインを+15.0上回る (55.0%→70.0%、相対+27.3%) - 結果の意義 - 進行的視覚予見 (progress-indexed visual foresight) が閉ループ制御に有効であることを実証

5. 議論はある?

- 要旨からは不明 - 具体的な議論や限界、今後の課題についての記述はない - 示唆される点 - 進行的視覚予見の価値が実証された - 効率性とOODロバスト性の両立が可能

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究 - World Action Models (WAMs) - 単一未来フレームを予測する最近のWAMs - 関連手法 - LIBERO-Plus, randomized RoboTwin, RoboCasa365 (ベンチマーク) - 同分野の定番 - ロボット制御のための視覚-言語-行動モデル (Vision-Language-Action models) - ビデオ予測に基づく計画手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar

分類: cs.AI, cs.CV, cs.RO

原文アブストラクト

World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in https://sii-ferenas.github.io/ProWAM-page.

関連論文

PR本紙発行元 EmplifAI