日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.40165

PrefPI: 選好ガイドによる分布外行動への誘導

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

シェア:XThreadsFacebookLINEはてブBluesky

自己生成軌道への相対選好のみを用い、事前学習済み生成ロボット方策を初期の有効範囲を超えた行動へ反復的に誘導する枠組みを提案。拡散方策やVLAで実機の物体運搬高さを10.7cmから19.8cmへと大幅に変化させた。

著者: Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajectories define a conditional distribution, whose density ratio with the broader behavior prior provides an implicit preference signal amplified by classifier-free guidance (CFG). Repeating this preference-conditioned modeling and guidance step yields a form of preference-guided policy iteration, turning incremental improvements toward previously inaccessible behaviors. Across diffusion policies and the PI0.5 flow- matching VLA in simulation and the real world, PrefPI produces substantial behavioral shifts with limited feedback. In particular, PrefPI increases object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.

関連論文

PR本紙発行元 EmplifAI