時間距離を考慮した表現による教師なし目標条件付き強化学習
TLDR: Unsupervised Goal-Conditioned RL via Temporal Distance-Aware Representations
時間距離を活用して探索と目標到達を促す教師なし目標条件付き強化学習手法TLDRを提案し、複数の移動環境で広い状態空間の獲得性能を向上させた。
著者: Junik Bae, Kwanyoung Park, Youngwoon Lee
分類: cs.LG, cs.AI
原文アブストラクト
Unsupervised goal-conditioned reinforcement learning (GCRL) is a promising paradigm for developing diverse robotic skills without external supervision. However, existing unsupervised GCRL methods often struggle to cover a wide range of states in complex environments due to their limited exploration and sparse or noisy rewards for GCRL. To overcome these challenges, we propose a novel unsupervised GCRL method that leverages TemporaL Distance-aware Representations (TLDR). Based on temporal distance, TLDR selects faraway goals to initiate exploration and computes intrinsic exploration rewards and goal-reaching rewards. Specifically, our exploration policy seeks states with large temporal distances (i.e. covering a large state space), while the goal-conditioned policy learns to minimize the temporal distance to the goal (i.e. reaching the goal). Our results in six simulated locomotion environments demonstrate that TLDR significantly outperforms prior unsupervised GCRL methods in achieving a wide range of states.