日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.06617

RealtimeWAM:ワンステップ非同期ワールドアクションモデル

RealtimeWAM: One-Step Asynchronous World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

映像生成バックボーンの表現を活用するワールドアクションモデルにおいて、教師アンカー型一貫性蒸留とクロスエキスパート・ウェーブフロント・パイプライン化により、ワンステップの行動生成と非同期推論を実現し、推論効率を大幅に向上させた研究。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) の一種で、video generation backbone の視覚表現を action 予測に使う枠組み。 - 既存の効率的 WAM は Mixture-of-Transformers (MoT) を採用し video 表現を一度計算して action expert が再利用する。 - しかし intra-expert iteration (multi-step action denoising) と inter-expert waiting (video/action expert の逐次実行) が推論効率を制限。 - RealtimeWAM は one-step action generation と asynchronous inference によりこの2つのボトルネックを解消する極めて効率的な WAM 変種。

2. 先行研究と比べてどこがすごい?

- 既存の効率的 WAM (Fast-WAM, Faster-WAM など) は video 表現を一度計算して再利用するが、multi-step action denoising と expert 間の逐次待ちが残る。 - RealtimeWAM は one-step action generation と asynchronous inference を導入し、これら2つのボトルネックを同時に解消。 - 多様な benchmark (LIBERO, LIBERO-Plus, RoboTwin) と model variant で near-lossless (<1% drop) を維持しつつ、H100 で約25倍の end-to-end 高速化を達成。

3. 技術・手法の肝は?

- Teacher-Anchored Consistency Distillation (TACD): local consistency error だけでは正確な最終 action を保証しない local-global error gap に対処。frozen teacher の multi-step rollout endpoint からの明示的 supervision を local consistency に補完し、正確な one-step action generation を可能にする。 - Cross-Expert Wavefront Pipelining (CEWP): expert 間の不要な待ちを排除。video KV cache を block 単位で共有し、対応する action attention が消費する直前にのみ同期することで2つの expert を overlap させる。

4. どうやって有効だと検証した?

- 多様な benchmark (LIBERO, LIBERO-Plus, RoboTwin) と model variant (Fast-WAM, Faster-WAM) で広範な実験を実施。 - RealtimeWAM が near-lossless 性能 (<1% drop) を維持しつつ、H100 で約25倍の end-to-end 高速化を達成することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Fast-WAM - Faster-WAM - Mixture-of-Transformers (MoT) - Teacher-Anchored Consistency Distillation (TACD) の基盤となる consistency distillation 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu, Shunzi Yang, Ruihao Gong, Shen Ren, Tianwei Zhang, Wenya Wang

分類: cs.CV, cs.LG, cs.RO

原文アブストラクト

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, $<1\%$ drop) across these benchmarks while delivering significant end-to-end speedup (\eg, $\sim25\times$ on H100). Our code and checkpoints are available via this \href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{link}.

関連論文

PR本紙発行元 EmplifAI