日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30247

Rolling-WAM: 転がる想像による世界行動モデル

Rolling-WAM: World Action Models with Rolling Imagination

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作のための世界行動モデルにおいて、映像と行動のノイズ除去を再計画サイクルごとに分割して行うことで、計算遅延を抑えつつ4.5倍の再計画高速化を実現した手法。

詳しい要約

1. どんなもの?

- ロボットマニピュレーションのためのWorld Action Models (WAMs)の新しい定式化であるRolling-WAMを提案。 - 従来のWAMsは各リプランニングサイクルでjoint video-action denoisingを完了するため遅延が大きい。 - Rolling-WAMはjoint denoisingを連続するリプランニングサイクルに分散させる。 - ビデオとアクションのチャンクを異なるノイズレベルでスライディングウィンドウに保持する。 - 各ステップでrolling noise scheduleにより、実行するアクションチャンクを完全にdenoiseし、将来のチャンクは部分的にrefineする。 - 新しいカメラ観測とともにウィンドウが進み、保持された将来チャンクのdenoisingが継続される。

2. 先行研究と比べてどこがすごい?

- 標準的なjoint WAMsと比較して、予測ホライズン全体をゼロからdenoiseする必要をなくす。 - これにより、定常状態でのリプランニング速度が4.5倍高速化される。 - 計算コストを時間的に分散させつつ、チャンク境界を越えて進化する視覚-行動コンテキストを保持する。 - LIBERO、RoboTwin、実世界のUnitree G1ヒューマノイドでの評価で、競争力のあるマニピュレーション性能を達成。

3. 技術・手法の肝は?

- ビデオとアクションのチャンクを異なるノイズレベルでスライディングウィンドウに維持する。 - rolling noise scheduleを導入し、各ステップで実行するアクションチャンクを完全にdenoiseし、将来のチャンクは部分的にrefineする。 - ウィンドウが新しいカメラ観測とともに進むにつれ、保持された将来チャンクのdenoisingプロセスが継続される。 - これにより計算コストが時間的に分散され、チャンク境界を越えた視覚-行動コンテキストが保持される。

4. どうやって有効だと検証した?

- LIBERO、RoboTwin、および実世界のUnitree G1ヒューマノイドでの評価を実施。 - Rolling-WAMが競争力のあるマニピュレーション性能を達成することを示した。 - 標準的なjoint WAMsと比較して、定常状態でのリプランニング速度が4.5倍高速化されることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 標準的なjoint WAMs。 - 関連手法: World Action Models (WAMs)、joint video-action denoising。 - データセット/ベンチマーク: LIBERO、RoboTwin、Unitree G1。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.

関連論文

PR本紙発行元 EmplifAI