PreferenceFlow: 人間介入によるフローマッチングロボット方策のテスト時ガイダンス
PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions
人間の介入とロボットの行動を比較して選好モデルを学習し、報酬なしで事前学習済みフローマッチング方策をテスト時に誘導する手法を提案。実機の精密挿入タスクで成功率を69%から90.5%に向上させた。
著者: Yiqi Tang, Diyuan Shi, Runze Li, Donglin Wang
分類: cs.RO
原文アブストラクト
Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a framework for improving a pretrained flow policy at test time without environment rewards or updates to the base policy. Human intervention chunks are paired with robot chunks generated from the same initial conditioning state to train a preference model. During infer- ence, we adopt the QGF sampling update, replacing its value gradient with the preference gradient evaluated at an estimated clean action. A gradient-cap loss penalizes excessive gradients on intervention pairs, while a zero-gradient loss discourages guidance near actions from expert demonstrations. On four real-world precision insertion tasks with a Franka robot, Pref- erenceFlow achieves a mean success rate of 90.5%, compared with 69% for the frozen policy. Ablations support the roles of gradient regularization and correctly ordered preference labels in the evaluated settings. These results demonstrate the utility of human interventions as local preference supervision for guiding frozen generative robot policies.