日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
スキル発見arXiv:2609.17682

拡散モデルによる多様で再利用可能な運動スキルの発見

DSD: Learning Diverse and Reusable Motor Skills via Diffusion Skill Discovery

シェア:XThreadsFacebookLINEはてブBluesky

拡散モデルで状態分布のエントロピー勾配を推定し、高次元ヒューマノイド制御において多様で再利用可能な運動スキルを獲得する手法を提案。

詳しい要約

1. どんなもの?

- 提案手法: Diffusion Skill Discovery (DSD) - 拡散モデルを用いてスキル発見を行う手法 - 高次元のヒューマノイド制御において多様で再利用可能な運動スキルを学習 - 目的: 下流タスクでの効率的な学習を可能にするスキルレパートリーの獲得 - 階層制御とゼロショット制御の2つの設定でスキルを再利用 - 特徴: スキル潜在変数と状態間の相互情報量最大化に基づく - 拡散モデルによるスコアマッチングで状態分布のエントロピー勾配を近似

2. 先行研究と比べてどこがすごい?

- 先行研究: 相互情報量最大化によるスキル発見 - 限界: 高次元制御では周辺状態エントロピーの直接推定が困難 - 従来は潜在空間近似や粗い状態分布推定に依存 - 提案手法の優位性: - 拡散モデルによるスコアマッチングでエントロピー勾配を正確に近似 - 状態空間の広いカバレッジを促進し、行動多様性を向上 - 下流タスクで再利用可能な複雑で敏捷な行動を発見

3. 技術・手法の肝は?

- 核心: 拡散モデルを利用したスキル発見 - ポリシーが生成する状態分布のエントロピー勾配をスコアマッチングで推定 - 目的関数: 周辺状態エントロピーと条件付きエントロピーの最大化 - 技術詳細: - 拡散モデルが状態分布のスコア(勾配)を学習 - スキル潜在変数と状態の相互情報量を最大化 - 高次元状態空間での効率的な探索を実現

4. どうやって有効だと検証した?

- 実験設定: 高次元ヒューマノイド制御タスク - 下流制御: 階層制御(タスク特化高レベルポリシー)とゼロショット制御(オフライン軌道からの潜在選択) - 結果: - 先行スキル発見手法よりも広範な再利用可能スキルを発見 - 複雑で敏捷な行動が下流タスクで再利用可能であることを確認

5. 議論はある?

- 議論点: - 拡散モデルの計算コストやサンプリング効率 - 他の高次元制御タスクへの汎用性 - スキルの解釈性や階層制御との統合 - 要旨からは不明: 具体的な限界や失敗ケース、計算資源の詳細

6. 次に読むべき論文は?

- 関連研究: - 相互情報量最大化に基づくスキル発見手法(DIAYN, DADSなど) - 拡散モデルを強化学習に応用した研究(Diffusion Policy, Decision Diffuserなど) - 階層的強化学習(HRL)やゼロショット制御 - 次に読むべき: これらの手法の原論文や拡散モデルとスキル発見の融合に関する最新研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sun Woo Kim, Xue Bin Peng

分類: cs.LG, cs.GR

原文アブストラクト

Humans efficiently learn new tasks by reusing a rich repertoire of motor skills across different goals and contexts. A similar strategy can also be used to enable simulated characters to efficiently perform new tasks by leveraging reusable motor skills. To support a wide range of downstream tasks, the learned repertoire should be diverse, consisting of distinct behaviors as well as spatial and temporal variation within each behavior. A commonly used method for learning diverse skills is by maximizing the mutual information between skill latents and the states produced by a policy. The marginal state entropy promotes broad behavioral coverage, while the conditional entropy encourages consistent behaviors from each latent. However, directly estimating the marginal state entropy is intractable in high-dimensional control problems. Prior methods therefore rely on indirect latent-space approximations or coarse estimators of the state distribution. These approximations may not effectively promote broad coverage of the state space, resulting in skills with limited behavioral diversity and reduced utility for downstream tasks. In this work, we propose Diffusion Skill Discovery (DSD), a skill discovery method that uses a diffusion model to approximate the entropy gradient of the policy-induced state distribution through score matching. The resulting objective encourages the discovery of skills that produce a broader range of behaviors for high-dimensional humanoid control. The learned skills are reused in two downstream control settings: hierarchical control with a task-specific high-level policy and zero-shot control through latent selection from offline trajectories. Our experiments show that DSD discovers a broader repertoire of reusable motor skills than prior skill discovery methods, leading to the emergence of complex and agile behaviors that can be reused across downstream tasks.

関連論文

PR本紙発行元 EmplifAI