数秒のデモから学ぶ四足歩行
Learning Quadruped Walking from Seconds of Demonstration
四足歩行の模倣学習が少量データで効果的な理由を解析し、潜在空間と出力の変動を整える新手法を提案。数秒のデモからオフラインで歩行ポリシーを学習できることを実機で確認した。
著者: Ruipeng Zhang, Hongzhan Yu, Ya-Chien Chang, Chenghao Li, Henrik I. Christensen, Sicun Gao
分類: cs.LG, cs.AI
原文アブストラクト
Quadruped locomotion provides a natural setting for understanding when model-free learning can outperform model-based control design, by exploiting data patterns to bypass the difficulty of optimizing over discrete contacts and the combinatorial explosion of mode changes. We give a principled analysis of why imitation learning with quadrupeds can be inherently effective in a small data regime, based on the structure of its limit cycles, Poincaré return maps, and local numerical properties of neural networks. The understanding motivates a new imitation learning method that regulates the alignment between variations in a latent space and those over the output actions. Hardware experiments confirm that a few seconds of demonstration is sufficient to train various locomotion policies from scratch entirely offline with reasonable robustness.