目標条件付き双模倣による転移可能なスキルの学習
Learning Transferable Skills using Goal-Conditioned Bisimulation
報酬なしデータから汎用方策を事前学習するため、環境の時間構造を捉えつつレイアウト変化に頑健な行動認識的時間表現を学習し、双模倣に基づく教師なしスキル発見により転移可能なスキルを獲得する手法を提案。
著者: Mohammad Amin Abbasfar, Farbod Azimmohseni, Mohammad Hossein Rohban
分類: cs.LG, cs.AI, cs.RO
原文アブストラクト
Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously unseen layouts. A key challenge is to learn representations that capture the temporal structure of the environment while remaining robust to variations across layouts. To address this issue, we present an objective for learning action-aware temporal representations that satisfy the functional equivariance property while preserving the local temporal structure of the environment. Building upon this embedding, we further propose unsupervised skill discovery using bisimulation, which learns transferable skills by conditioning the behavior of skills exclusively on the subset of state features that directly affect their execution. This enforces invariant behavior across different layouts, enabling skills to transfer effectively to other configurations. Finally, through comprehensive empirical evaluations, we show that skills learned in a given environment can be effectively applied to solve downstream tasks in various environment layouts, demonstrating strong out-of-distribution generalization.