日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24868

DualWAM: 非同期グローバル計画と局所修正のためのデュアルシステム世界行動モデル

DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement

シェア:XThreadsFacebookLINEはてブBluesky

グローバル計画と手首視点による局所修正を非同期に分離し、高頻度な閉ループ制御を可能にした世界行動モデル。FrankaとGalbotでのゼロショット操作で成功率を平均4.5ポイント向上させ、16.6倍の高速化を実現。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) はロボットの行動生成と将来の世界状態予測を同時に行う枠組み。 - 将来の視覚予測は計算コストが高く、既存 WAMs は長い action chunk で推論コストを償却するため閉ループ応答性が犠牲になる。 - 本研究は DualWAM を提案。global planning と local refinement を分離した dual-system WAM。 - 広い時間軸の world-action 生成を保ちつつ、高頻度の閉ループ行動更新を可能にする。

2. 先行研究と比べてどこがすごい?

- 既存 WAMs は長い action chunk に依存し、閉ループ応答性が低いという課題があった。 - DualWAM は global planning と local refinement を非同期に分離し、広い時間軸の生成を維持しつつ高頻度更新を実現。 - Franka と Galbot での zero-shot 操作タスクで、最強の評価済み baseline より平均 4.5 ポイント成功率が向上。 - さらに critical-path で 16.6 倍の高速化を達成。

3. 技術・手法の肝は?

- 二つのシステムから構成。 - グローバル側 (system two) は、より広い world-action chunk に対して高ノイズの双方向 denoising を定期的に実行し、global plan を確立。 - ローカル側 (system one) は wrist-only で、中間 denoising 状態から時間的に整合した短い window を抽出し、最新の wrist 観測を用いて低ノイズ refinement を完了。 - wrist 観測は相互作用中の局所的な geometry、motion、contact に関する action-aligned な手がかりを提供。 - 両システムは共有された denoising trajectory に沿って非同期に動作。各 global plan は複数の local update で再利用され、system one は新鮮な interaction feedback を繰り返し取り込む。

4. どうやって有効だと検証した?

- Franka と Galbot での zero-shot 操作タスクで評価。 - 最強の評価済み baseline より平均 4.5 ポイント成功率が向上。 - critical-path で 16.6 倍の高速化を達成。 - 追加研究で、役割に合った egocentric データと UMI データが成功率を 14 ポイント改善することを示す。 - 分離設計が edge-cloud 展開を自然にサポートし、baseline より通信オーバーヘッドが大幅に低いことも示す。

5. 議論はある?

- 役割に合った egocentric および UMI データが成功率を 14 ポイント改善。 - 分離設計は edge-cloud 展開を自然にサポートし、baseline より通信オーバーヘッドが大幅に低い。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: World Action Models (WAMs)、baseline として評価された最強手法。 - 関連手法: egocentric データ、UMI データを用いた学習。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yixin Zheng, Jiangran Lyu, Yuntian Deng, Kai Liu, Yizhou Zhou, Yizhou Wang, Xiaoguang Zhao, He Wang, Zhizheng Zhang

分類: cs.RO

原文アブストラクト

World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at the cost of closed-loop responsiveness. We present \method, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement. \systwo periodically performs high-noise bidirectional denoising over a broader world-action chunk to establish a global plan, while wrist-only \sysone extracts a temporally aligned short window from the intermediate denoising state and completes low-noise refinement using the latest wrist observations, which provide action-aligned cues about local geometry, motion, and contact during interaction. The two systems operate asynchronously along a shared denoising trajectory: each global plan is reused across multiple local updates, while \sysone repeatedly incorporates fresh interaction feedback. Across zero-shot manipulation tasks on Franka and Galbot, \method improves success over the strongest evaluated baseline by 4.5 percentage points on average, while achieving a 16.6$\times$ critical-path speedup. Further studies show that role-matched egocentric and UMI data improve success by 14 percentage points, and that the decoupled design naturally supports edge--cloud deployment with substantially lower communication overhead than the baseline.

関連論文

PR本紙発行元 EmplifAI