日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2607.23726

疎な報酬・長期的課題のための階層型ソフトアクター批判

Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

疎な報酬と長い時間軸を持つ強化学習タスクに対し、高レベル戦略と低レベル連続制御を組み合わせた階層型強化学習フレームワークを提案し、SAR-2データセットで有効性を実証した。

著者: Zahra Abdalla Elashaal, Afef Hfaiedh, Nahla Khraief, Issmail Ellabib, Giansalvo Cirrincione

分類: cs.RO, cs.LG

原文アブストラクト

Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.

関連論文