日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン強化学習arXiv:2405.14790

DIDI: 拡散モデル誘導によるオフライン行動生成の多様性

DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

シェア:XThreadsFacebookLINEはてブBluesky

拡散確率モデルを事前分布として用い、ラベルなしオフラインデータから多様なスキルを学習する手法を提案。

著者: Jinxin Liu, Xinghong Guo, Zifeng Zhuang, Donglin Wang

分類: cs.LG

原文アブストラクト

In this paper, we propose a novel approach called DIffusion-guided DIversity (DIDI) for offline behavioral generation. The goal of DIDI is to learn a diverse set of skills from a mixture of label-free offline data. We achieve this by leveraging diffusion probabilistic models as priors to guide the learning process and regularize the policy. By optimizing a joint objective that incorporates diversity and diffusion-guided regularization, we encourage the emergence of diverse behaviors while maintaining the similarity to the offline data. Experimental results in four decision-making domains (Push, Kitchen, Humanoid, and D4RL tasks) show that DIDI is effective in discovering diverse and discriminative skills. We also introduce skill stitching and skill interpolation, which highlight the generalist nature of the learned skill space. Further, by incorporating an extrinsic reward function, DIDI enables reward-guided behavior generation, facilitating the learning of diverse and optimal behaviors from sub-optimal data.

関連論文

PR本紙発行元 EmplifAI