日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34608

適応的中间状態による効率的なWorld Action Model推論

Efficient World Action Model Inference with Adaptive Intermediate States

シェア:XThreadsFacebookLINEはてブBluesky

World Action Modelの推論を、軌道リマッピング・観測リバインディング・残差リスケーリングで高速化する学習不要フレームワークを提案。

著者: Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, Ting Cao, Yunxin Liu, Kun Li

分類: cs.RO, cs.AI

原文アブストラクト

World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.

関連論文

PR本紙発行元 EmplifAI