PF-RL: 視覚言語行動モデルのための目標条件付き価値幾何による進捗場強化学習
PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models
VLAモデルの特徴上に目標条件付きの進捗表現を学習し、その幾何的距離から密な報酬を生成することで、長期的なマニピュレーションタスクの強化学習を効率化する手法を提案。
著者: Yunpeng Qing, Yilun Kong, Sixu Lin, Ming Zhou, Yiming Fei, Shuang Luo, Yixiao Chi, Haoming Gu, Jingyuan Liu, Changxu Wei, Zhi Hou, Changqing Zou
分類: cs.RO
原文アブストラクト
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.