日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11401

WAM-Cache: 古さを制限したKV再利用による効率的なWorld Action Model

WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

World Action Modelの動画DiTによるKVキャッシュをチャンク間で再利用し、行動エキスパートの注意と視覚的驚きに基づくトークンだけを再計算することで、精度を保ちつつ計算量を削減する学習不要の手法。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) の closed-loop control を対象とした、training-free な推論高速化フレームワーク。 - 各 chunk で video Diffusion Transformer (DiT) が観測を layerwise key-value (KV) に符号化する prefill が計算コストを支配する点に着目。 - layerwise KV を chunk 間で保持し、sparse refresh set の token のみ再計算する。 - Fast-WAM 上で video DiT prefill FLOPs を 32-42% 削減。

2. 先行研究と比べてどこがすごい?

- 既存の training-free acceleration は prefill を fully dense のまま扱っていた。 - WAM-Cache は KV を chunk 間で再利用し、sparse refresh のみ行う点で異なる。 - 直感的な『視覚的に drift した token を refresh』という heuristic は、oracle で真の KV drift を予測しても dense baseline を大きく下回ることを発見。 - 代わりに action expert の cross-attention を基準に refresh set を選ぶことで dense に近い精度を維持。

3. 技術・手法の肝は?

- layerwise KV を chunk 間で保持し、sparse refresh set の token のみ再計算。 - refresh set は action expert の cross-attention と visual latent surprise を統合して選択。 - さらに strict age bound を導入し、誤差の累積を抑制。 - 学習不要 (training-free) で既存 WAM に適用可能。

4. どうやって有効だと検証した?

- Fast-WAM 上で RoboTwin 2.0、LIBERO、実世界実験を実施。 - video DiT prefill FLOPs を 32-42% 削減。 - シミュレーションで dense policy との差は 0.7-1.8 percentage points 以内。 - 実機で 2.5 points 以内。

5. 議論はある?

- 視覚的 drift に基づく refresh は oracle でも不十分であり、action expert の attention が重要と指摘。 - strict age bound が誤差累積抑制に寄与。 - 精度低下はシミュレーションで 0.7-1.8 points、実機で 2.5 points と報告。 - 他の WAM やタスクへの一般性は要旨からは不明。

6. 次に読むべき論文は?

- Fast-WAM (本手法の適用先) - video Diffusion Transformer (DiT) - World Action Models (WAMs) - RoboTwin 2.0 - LIBERO

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kai Ding, Yang He, Ruijie Quan, Yi Yang

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.

関連論文

PR本紙発行元 EmplifAI