日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2403.09930

品質多様性アクタークリティック:価値と後継特徴クリティックによる高性能で多様な行動の学習

Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics

シェア:XThreadsFacebookLINEはてブBluesky

価値関数と後継特徴の2つのクリティックを制約付き最適化で統合し、高いリターンと多様なスキルを同時に獲得するオフポリシー強化学習手法QDACを提案した。

著者: Luca Grillotti, Maxence Faldor, Borja G. León, Antoine Cully

分類: cs.LG, cs.AI

原文アブストラクト

A key aspect of intelligence is the ability to demonstrate a broad spectrum of behaviors for adapting to unexpected situations. Over the past decade, advancements in deep reinforcement learning have led to groundbreaking achievements to solve complex continuous control tasks. However, most approaches return only one solution specialized for a specific problem. We introduce Quality-Diversity Actor-Critic (QDAC), an off-policy actor-critic deep reinforcement learning algorithm that leverages a value function critic and a successor features critic to learn high-performing and diverse behaviors. In this framework, the actor optimizes an objective that seamlessly unifies both critics using constrained optimization to (1) maximize return, while (2) executing diverse skills. Compared with other Quality-Diversity methods, QDAC achieves significantly higher performance and more diverse behaviors on six challenging continuous control locomotion tasks. We also demonstrate that we can harness the learned skills to adapt better than other baselines to five perturbed environments. Finally, qualitative analyses showcase a range of remarkable behaviors: adaptive-intelligent-robotics.github.io/QDAC.

関連論文

PR本紙発行元 EmplifAI