日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
逆強化学習arXiv:2609.31855

PHIRL: 逆強化学習のための学習報酬とタスク進捗の整合

PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

人間のデモンストレーションと進捗フィードバックを組み合わせ、逆強化学習でロバストな報酬関数を学習するデータ効率の良いフレームワークPHIRLを提案。実機・シミュレーションでベースラインを上回る性能を示した。

著者: Hang Yu, James Staley, Cheng Xi Tsou, Xiujin Liu, Wenchang Gao, Jindan Huang, Shijie Fang, Zhegong Shangguan, Angelo Cangelosi, Reuben Aronson, Elaine Short

分類: cs.RO, cs.AI

原文アブストラクト

Human demonstrations provide dense policy-level information but sometimes lack local precision. Human feedback presents accurate local critiques, but offers sparse evaluations rather than direct policy guidance. We propose Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that learns robust reward functions by jointly leveraging demonstrations and feedback. Specifically, we use progress, a feedback modality that describes cumulative task completion. PHIRL iteratively infers a reward function from demonstrations via inverse reinforcement learning, calculates the learned rewards over the progress-annotated demonstrations, and aligns the rewards with progress annotations over four dimensions. We evaluate PHIRL on real and simulated robot tasks, with additional exploration using a fine-tuned vision-language model to provide progress feedback. Results demonstrate that PHIRL significantly outperforms the baselines, achieving substantially higher environmental return rewards and task success with only twenty percent of demonstrations annotated. Analysis of reward-hacking scenarios demonstrates that PHIRL learned reward functions are reliable against exploitation.

関連論文

PR本紙発行元 EmplifAI