日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.10601

ロボット犬に新たな芸を教える:敵対的タスク選択を用いた強化学習と模倣学習の組み合わせによる多様な四足歩行スキル

Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

シェア:XThreadsFacebookLINEはてブBluesky

強化学習と模倣学習を組み合わせ、敵対的タスク選択で複数の教師方策から生徒方策を訓練し、四足ロボットに歩行・掘削・跳躍など22の多様なスキルを習得させる手法を提案した。

著者: Lemon Foxmere, Anthony Furman, Yizheng Du, Oliver Chang, Leilani Gilpin, Steve McGuire

分類: cs.RO

原文アブストラクト

Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional stages train a student policy with a multi-teacher distillation setup that uses a combined RL and Imitation Learning (IL) objective under an adversarial task selection process that focuses training on the worst-performing task. With this method, we train a student policy that performs 22 tasks using 8 teachers. Evaluations show our method preserves motion quality and tracks commands more accurately than PPO and distill-then-finetune baselines, and in some cases generalizes to new tasks without explicit training. Finally, we demonstrate real-world robustness by deploying the resulting policy on a Unitree B1 quadruped. Video: https://youtu.be/V9yX04EBcFA

関連論文

PR本紙発行元 EmplifAI