日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.32155

RecastVLA: 適応的ポリシー状態による過去の相互作用から未来の制御へ

RecastVLA: From Past Interaction to Future Control with Adaptive Policy States

シェア:XThreadsFacebookLINEはてブBluesky

フローマッチング型VLAポリシー内に適応的ポリシー状態を保持し、テスト時訓練で過去の行動生成履歴を次の制御に活かす手法を提案。LIBEROやRoboTwin、実機タスクで成功率を向上させた。

著者: Wenbo Li, Jun Yang, Yiteng Chen, Wei Zhang, Qingyao Wu

分類: cs.RO

原文アブストラクト

Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read the same state, while features across depths and flow evaluations jointly define one update for the next policy call. Subsequent action losses train the initialization, interfaces, and update rule by differentiating through earlier state transitions. At deployment, updates use the policy's own action-generation features without expert action labels. Across LIBERO, RoboTwin, RoboDojo, and twelve real-robot tasks, RecastVLA improves mean success over a matched policy trained without test-time training, including 10.68 percentage points on RoboTwin Clean-to-Clean. In controlled RoboTwin comparisons, retaining state improves success, and the shared design exceeds independently trained layer-local TTT by 2.58 points.

関連論文

PR本紙発行元 EmplifAI