日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10528

Long-WAM: 世界行動モデルのコンテキスト拡張

Long-WAM: Scaling the Context of World-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

因果的世界行動モデルのコンテキストを拡張し、自己回帰的事前学習とストリーミング符号化によりリアルタイム制御を実現したフレームワーク。

詳しい要約

1. どんなもの?

- リアルタイムなロボット制御のためのモデル・システムフレームワーク - causal world-action model の context を拡張する Long-WAM を提案 - 視覚履歴を活用して動作やタスク進捗を推論 - 履歴アクセスと履歴利用は異なるという中心知見 - 映像基盤を autoregressive (AR) で事前学習すると長い履歴が有効 - ロボットおよび egocentric 動画から action label なしで causal prediction を学習 - world-action adaptation 中に history-to-future 構造を保持 - RTX 5090, DGX Spark, Jetson AGX Thor 上で展開可能 - Unitree G1 と YAM でのリアルタイム展開を実現

2. 先行研究と比べてどこがすごい?

- 従来の bidirectionally pretrained 初期化では context 拡張の利得がほぼない - AR 事前学習により context を 0.0 秒から 19.2 秒に増やすと成功率が 63.3% から 78.7% に向上 - robot-domain AR pretraining が GR-1 と LIBERO-Long のピーク成功率をさらに向上 - LIBERO-Long, RoboTwin 2.0, DOMINO で比較手法中最高性能 - 動的カップ積みで 95% 成功、Pi0.5 と Fast-WAM は 20 試行中 0 成功 - メモリを活用した実行者として高レベル計画を補完

3. 技術・手法の肝は?

- causal world-action model の context をスケールするモデル・システムフレームワーク - ロボットおよび egocentric 動画から action label なしで causal prediction を学習 - world-action adaptation 中に history-to-future 構造を保持 - streaming observation encoding - asynchronous execution - hardware-specific acceleration - 未来映像潜在予測を含む action chunk を 107.4 ms で処理 (RTX 5090) - リアルタイム制約下で context を拡張

4. どうやって有効だと検証した?

- RoboCasa GR-1 で context 0.0→19.2 秒で成功率 63.3%→78.7% - bidirectionally pretrained 初期化では正味の利得なし - robot-domain AR pretraining で GR-1 と LIBERO-Long のピーク成功率向上 - LIBERO-Long, RoboTwin 2.0, DOMINO で比較手法中最高性能 - RTX 5090, DGX Spark, Jetson AGX Thor で展開検証 - Unitree G1 と YAM で動的・長地平操作を実証 - 動的カップ積みで 95% 成功、Pi0.5 と Fast-WAM は 20 試行中 0 成功

5. 議論はある?

- 履歴へのアクセスと履歴の利用は異なる - 長い履歴の利得は AR 事前学習に強く依存 - bidirectional 初期化では context 拡張の効果が限定的 - リアルタイム制約下での context 拡張のトレードオフ - メモリを活用した実行者として高レベル計画を補完 - 詳細な限界や失敗事例は要旨からは不明

6. 次に読むべき論文は?

- Pi0.5 - Fast-WAM - RoboCasa GR-1 - LIBERO-Long - RoboTwin 2.0 - DOMINO - Unitree G1 - YAM - causal world-action model - autoregressive video foundation model

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu, Shuai Yang, Weian Mao, Luozhou Wang, Yicheng Xiao, Weifeng Lin, Qixin Hu, Bryan Chu, Sifei Liu, Linxi Fan, Xiaojuan Qi, Song Han, Yukang Chen

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.

関連論文

PR本紙発行元 EmplifAI