日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2604.24127

人間フィードバックを活用した意味的に有用なスキル発見

Leveraging Human Feedback for Semantically-Relevant Skill Discovery

シェア:XThreadsFacebookLINEはてブBluesky

強化学習における教師なしスキル発見に人間のフィードバックを組み込み、意味的に多様で関連性のあるスキルを効率的に発見する手法を提案した。

著者: Maxence Hussonnois, Thommen George Karimpanal, Santu Rana

分類: cs.LG, cs.AI

原文アブストラクト

Unsupervised skill discovery in reinforcement learning aims to intrinsically motivate agents to discover diverse and useful behaviours. However, unconstrained approaches can produce unsafe, unethical, or misaligned behaviours. To mitigate these risks and improve the practical desireability of discovered skills, recent work grounds the discovery process by leveraging human preference feedback. However, preference-based approaches are feedback-inefficient and inherently ill-equipped to deal with skill spaces composed of a variety of different skills such as running, jumping, walking, etc. To overcome this limitation, we introduce semantic labelling, a novel and feedback-efficient approach that leverages human cognitive strengths to identify and label semantically meaningful behaviours. Based on semantic labelling, we propose Semantically Relevant Skill Discovery (SRSD), a novel human-in-the-loop approach that collects semantic labels from human feedback and learns a reward function to encourage skills to be more semantically diverse and relevant. Through our experiments in a 2D navigation environment and four locomotion environments, we demonstrate that SRSD can improve semantic diversity and discover relevant behaviours while scaling effectively to a large variety of behaviours.

関連論文