振る舞い基盤モデルによるゼロショット全身ヒューマノイド制御
Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models
ラベルなし動作キャプチャデータから、報酬や模倣タスクにゼロショットで汎化できるヒューマノイドの振る舞い基盤モデルを学習する手法を提案。
著者: Andrea Tirinzoni, Ahmed Touati, Jesse Farebrother, Mateusz Guzek, Anssi Kanervisto, Yingchen Xu, Alessandro Lazaric, Matteo Pirotta
分類: cs.LG
原文アブストラクト
Unsupervised reinforcement learning (RL) aims at pre-training agents that can solve a wide range of downstream tasks in complex environments. Despite recent advancements, existing approaches suffer from several limitations: they may require running an RL process on each downstream task to achieve a satisfactory performance, they may need access to datasets with good coverage or well-curated task-specific samples, or they may pre-train policies with unsupervised losses that are poorly correlated with the downstream tasks of interest. In this paper, we introduce a novel algorithm regularizing unsupervised RL towards imitating trajectories from unlabeled behavior datasets. The key technical novelty of our method, called Forward-Backward Representations with Conditional-Policy Regularization, is to train forward-backward representations to embed the unlabeled trajectories to the same latent space used to represent states, rewards, and policies, and use a latent-conditional discriminator to encourage policies to ``cover'' the states in the unlabeled behavior dataset. As a result, we can learn policies that are well aligned with the behaviors in the dataset, while retaining zero-shot generalization capabilities for reward-based and imitation tasks. We demonstrate the effectiveness of this new approach in a challenging humanoid control problem: leveraging observation-only motion capture datasets, we train Meta Motivo, the first humanoid behavioral foundation model that can be prompted to solve a variety of whole-body tasks, including motion tracking, goal reaching, and reward optimization. The resulting model is capable of expressing human-like behaviors and it achieves competitive performance with task-specific methods while outperforming state-of-the-art unsupervised RL and model-based baselines.
関連論文
- ResSafe: 残差強化学習によるヒューマノイドの安全フィルタリングヒューマノイド制御
- ViBe: 知覚型ヒューマノイド全身制御のための視覚行動適応ヒューマノイド制御
- GLoRI: グローバル・ローカル参照相互作用によるヒューマノイド移動操作のための閉ループ全身追跡ヒューマノイド制御
- 学習された停止可能性値によるヒューマノイドの安全停止ヒューマノイド制御
- ADAPT: 俊敏な拡散行動事前分布による堅牢で操縦可能なオンライン文章駆動ヒューマノイド制御ヒューマノイド制御
- StableMimic: 人間らしいスムーズな復帰を実現するヒューマノイド動作追跡 - 追跡分布を超えた学習による構造化された転倒後行動ヒューマノイド制御