日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35690

エージェント事前分布に導かれた政策学習

Agent Priors-guided Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

各スキルの構造的事前分布を政策のインターフェースとして用いることで、スキルの汎化と新規タスクへの組み合わせを両立させる手法APPLを提案。

詳しい要約

1. どんなもの?

ロボットが少数のdemonstrationから学ぶ際に必要となる、compositional generalizationとskill generalizationの両立を目指す手法。各skillのpolicyが持つstructural priorを、compositionとskillのinterfaceとして用いる。これを実装したAgent Priors-guided Policy Learning (APPL)を提案。construction agentがdemonstrationをskillに分割し、各skillに複数のstructural priorを提案、priorごとにpolicyを訓練・検証する。runtime agentがprior-specific policyを選択し、interfaceを用いて新task goalへcomposeする。

2. 先行研究と比べてどこがすごい?

従来はcompositionがskillをname・instruction・symbolic operator等の別記述を通してのみ参照し、policyが訓練された構造の情報が失われていた。本研究は各policyのstructural priorをinterfaceの一部として明示的に用いる点が新しい。これによりcompositionがpolicyの適用範囲を把握でき、out-of-distributionなskill generalizationの改善と未見のskill compositionを可能にする。

3. 技術・手法の肝は?

鍵はstructural priorの活用。structural priorはbehaviorが何に依存するかを述べるもの(例: graspはgripperのobjectに対するposeのみに依存)。これを訓練に組み込むことでpolicyのgeneralize範囲を形成し、言語で記述することでcompositionに適用範囲を伝える。APPLではconstruction agentが完全なdemonstrationを再利用可能なskillにsegmentし、各skillに複数のstructural priorを提案、priorごとにpolicyを訓練・検証する。runtime agentはprior-specific policyを選択し、interfaceを用いて新goalへcomposeする。

4. どうやって有効だと検証した?

MetaWorldとlong-horizon ManiSkillタスクで評価。APPLはout-of-distributionなskill generalizationを改善し、未見のskill compositionを可能にした。interface情報をablateすると性能が大幅に低下した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法としてMetaWorldやManiSkillを用いたskill learning、compositional generalization、policy learningの研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Puming Jiang, Tianrun Hu, Haozhe Du, Yibo Li, Zhiwei Xue, Xinhu Li, Harold Soh

分類: cs.RO

原文アブストラクト

Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.

関連論文

PR本紙発行元 EmplifAI