日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.07922

OpenWAM: 構成可能なWorld-Actionモデルのためのオープンフレームワーク

OpenWAM: An Open Framework for Composable World-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

ロボット動画の基盤モデルと構成可能な映像-行動相互作用を組み合わせたオープンなWorld-Actionモデリング枠組みを提案し、LIBEROや実機タスクで高い成功率を達成した。

詳しい要約

1. どんなもの?

- OPENWAMは、world-action models (WAMs) のためのオープンなフレームワーク。 - 共通のcausal robot-video foundationと、設定可能なvideo-action interactionを中心に構築。 - Wan2.2-5Bを起点に、10,000時間以上の動画でcausal robot-video pretrainingを実施。 - action expertをshared Mixture-of-Transformersアーキテクチャで統合し、joint、video-then-action、action-then-video、decoupled generationをサポート。 - ロボット制御と未来予測を結合するWAMの設計選択を比較可能にする。

2. 先行研究と比べてどこがすごい?

- 既存のWAMはvideo backbone、interaction structure、supervision、inference procedureを同時に変化させており、設計選択の比較が困難だった。 - OPENWAMは共通の基盤と設定可能な相互作用構造を提供し、公平な比較を可能にする。 - causal robot-video training with causal adaptationにより、LIBERO-LongでのVTA successが68.4%から97.8%に向上。 - 同じアーキテクチャがinverse dynamicsとforward dynamicsに自然に拡張され、counterfactual transitionsの効果を研究できる。

3. 技術・手法の肝は?

- Wan2.2-5Bをベースに、10,000時間以上の動画でcausal robot-video pretrainingを実施。 - shared Mixture-of-Transformersアーキテクチャでaction expertを統合。 - joint、video-then-action、action-then-video、decoupled generationの4つの相互作用モードを設定可能。 - inverse dynamicsとforward dynamicsにも同じ設定可能なアーキテクチャを拡張。 - counterfactual transitionsを活用し、デモンストレーションを超えたダイナミクス学習を研究。

4. どうやって有効だと検証した?

- 4つのLIBEROスイートと実世界のbimanualタスクで高い成功率を達成。 - LIBERO-LongでのVTA successが68.4%から97.8%に改善。 - 新しいタスクにvideo predictorのみを適応した場合、counterfactual dataとdemonstrationsで訓練したfrozen local-context inverse dynamics modelが、4つのheld-out LIBERO-90タスクで平均84.0%の成功率を達成。 - 比較として、full-context inverse modelは47.0%、demonstrationsのみで訓練したlocal-context modelは21.5%。 - forward dynamicsでは、counterfactual supervisionによりRGB prediction errorが34.5%減少し、16のsame-state outcomes間のoutcome identificationが21.1%から71.3%に向上。

5. 議論はある?

- WAMの相互作用設計を比較するための共通テストベッドを提供。 - 成功したデモンストレーションを超えたvideo dataからのダイナミクス学習を研究可能。 - counterfactual transitionsが独立に訓練されたダイナミクスモデルを改善することを示唆。 - 限界や課題については要旨からは不明。

6. 次に読むべき論文は?

- Wan2.2-5B(ベースモデル) - LIBERO(ベンチマーク) - LIBERO-Long、LIBERO-90(評価タスク) - Mixture-of-Transformers(アーキテクチャ) - inverse dynamics model、forward dynamics model(関連手法) - counterfactual transitions(関連概念)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli

分類: cs.RO, cs.CV

原文アブストラクト

World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.

関連論文

PR本紙発行元 EmplifAI