想像した通りに実現する:視覚プランと行動の整合学習
Achieve What You Imagined: Learning to Align Actions with Visual Plans
世界行動モデルが生成する視覚的予測を目標条件付き提案として扱い、凍結した行動条件付き世界モデルとの予測一致度をフィードバックにFlow Policy Optimizationで行動ヘッドを最適化する手法を提案。実機UR5の4タスクで成功率を43.4%から75.1%に向上させた。
著者: Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou, Sichao Liu
分類: cs.RO, cs.AI, cs.LG
原文アブストラクト
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $π_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/