日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
選好ベース強化学習arXiv:2409.07268

多種類選好学習:等価選好を活用した選好ベース強化学習

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

シェア:XThreadsFacebookLINEはてブBluesky

人間の等価選好も学習に取り入れることで、選好ベース強化学習のフィードバック効率を向上させる手法を提案し、複数のタスクで有効性を検証した。

著者: Ziang Liu, Junjie Xu, Xingjiao Wu, Jing Yang, Liang He

分類: cs.LG

原文アブストラクト

Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teachers may choose equal preferences. This neglect may hinder the understanding of the agent regarding the task perspective of the teacher, leading to the loss of important information. To address this issue, we introduce the Equal Preference Learning Task, which optimizes the neural network by promoting similar reward predictions when the behaviors of two agents are labeled as equal preferences. Building on this task, we propose a novel PBRL method, Multi-Type Preference Learning (MTPL), which allows simultaneous learning from equal preferences while leveraging existing methods for learning from explicit preferences. To validate our approach, we design experiments applying MTPL to four existing state-of-the-art baselines across ten locomotion and robotic manipulation tasks in the DeepMind Control Suite. The experimental results indicate that simultaneous learning from both equal and explicit preferences enables the PBRL method to more comprehensively understand the feedback from teachers, thereby enhancing feedback efficiency. Project page: \url{https://github.com/FeiCuiLengMMbb/paper_MTPL}

関連論文

PR本紙発行元 EmplifAI