WAM-OPD: ワールドアクションモデルのためのオンポリシー蒸留
WAM-OPD: On-Policy Distillation for World Action Models
ビデオ生成とロボット動作生成を統合したワールドアクションモデル(WAM)の蒸留後性能低下を、オンポリシー蒸留で修復する手法を提案。学生モデルが環境で行動し、教師モデルがその履歴に基づいてビデオと動作のターゲットを提供することで、タスク成功率が大幅に向上した。
著者: Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang
分類: cs.AI, cs.RO
原文アブストラクト
World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. We study whether on-policy distillation (OPD) can repair such a student without requiring sparse-reward reinforcement learning. We introduce WAM-OPD, a deployment-consistent post-training recipe for a video-first WAM. The student acts in the environment and therefore determines the history distribution. A frozen teacher labels those student histories with coherent video and action targets, while the student action branch is trained under its own generated video plan, as it is at deployment. Joint video and action losses update lightweight adapters in the shared backbone, together with an action flow-matching regularizer. In preliminary RoboTwin 2.0 studies on two tasks, the released one-video/one-action-step Flash-WAM improves from 0.0% to 58.3% success on HANDOVER MIC, and from 16.7% to 33.3% on PUT OBJECT CABINET. These task-specific results are an initial capability proof rather than evidence of broad or uniform generalization. They nevertheless suggest that dense teacher supervision on student-induced histories is a promising post-training interface for video-first WAMs.