日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10079

RealtimeWAM: ワールドアクションモデルをどこまで高速に動かせるか

RealtimeWAM: How Fast Can I Run My World Action Model?

シェア:XThreadsFacebookLINEはてブBluesky

ワールドアクションモデルの推論を、学習不要の並列実行と適応的計算の協調により高速化する汎用フレームワークを提案し、複数ベンチマークで約9〜11倍の速度向上と成功率維持を実現した。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) の推論を高速化する汎用・training-free なフレームワーク RealtimeWAM を提案。 - visual dynamics modeling と action generation を組み合わせた WAMs は推論遅延が大きく、responsive な robot control を妨げる。 - 並列実行と adaptive computation を協調させ、多様な WAM アーキテクチャで低遅延推論を実現する。

2. 先行研究と比べてどこがすごい?

- FastWAM は test time の future-video generation を除去するが、専用の architectural design を要する。 - 一般的な caching 戦略は feature redundancy を利用するが、closed-loop control の変化する計算需要を捉えられない。 - RealtimeWAM は training-free かつ汎用的で、特定アーキテクチャに依存せず低遅延化できる点が異なる。

3. 技術・手法の肝は?

- layerwise dependencies を利用し、observation processing と prediction を overlap させる並列実行を行う。 - 並列分岐が GPU 資源を競合するため、pipeline 全体で adaptive computation を適用する。 - selective reuse として、視覚的に安定した領域の observation features を caching し、Transformer residuals を再利用する。 - 小さな predicted adjustments には追加の refinement を割り当てる。

4. どうやって有効だと検証した?

- FastWAM と OpenWAM を RoboTwin, LIBERO, LIBERO-Plus で評価。 - RTX 4090 上で mean inference latencies は 24.09 ms と 63.09 ms、平均 speedups は 8.90x と 10.67x。 - 平均 success rates は 82.75% と 87.41% で、native inference との差は 0.02 と 0.53 percentage points。 - 5 つの real-world tasks で FastWAM と OpenWAM の平均 success rates が native inference より 17.2 と 37.2 percentage points 向上。

5. 議論はある?

- 並列実行だけでは GPU 資源競合により利得が制限される点を指摘。 - adaptive computation により closed-loop control の変化する計算需要に対応する必要性を議論。 - 精度低下は native inference 比で僅少に抑えられている。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- FastWAM - OpenWAM - RoboTwin - LIBERO - LIBERO-Plus

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huanan Liu, Ye Li, Kangye Ji, Xiaoyu Chen, Hanyun Cui, Yutian Shen, Yuan Meng, Chenglei Wu, Jingyan Jiang, Bo Li, Zhi Wang

分類: cs.RO

原文アブストラクト

World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90$\times$ and 10.67$\times$. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.

関連論文

PR本紙発行元 EmplifAI