日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.10489

HuMBLE: 人間の動作データ駆動による身体性ロコモーションの行動学習

HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion

シェア:XThreadsFacebookLINEはてブBluesky

人間の動作データから自然でロバストな二足歩行ポリシーを学習するフレームワークを提案し、AtlasやUnitree G1で実機検証した。

詳しい要約

1. どんなもの?

- 人間の動作データから、リアルタイムで操縦可能かつ堅牢で生体模倣的な歩行ポリシーを学習するフレームワーク。 - 人間の歩行データセットを用い、教師-生徒蒸留で自然な歩行の事前ポリシーを学習。 - その後、マルチタスクRLで微調整し、コマンド追従性とロバスト性を向上。 - 3体の人型ロボット(Atlas R1, Atlas D1, Unitree G1)で検証。

2. 先行研究と比べてどこがすごい?

- 従来のコマンド追従・ロバスト性重視のRLポリシーは機械的な歩行になりがち。 - 人間動作データに依存するポリシーはデータ分布外のコマンドに汎化しにくい。 - 本手法は両者のバランスを取り、人間らしさを保ちつつ任意のコマンドに追従可能。 - 人間データなしで学習したTabula Rasa RLポリシーとの比較で優位性を確認。

3. 技術・手法の肝は?

- 全身参照条件付きRLポリシーを学習し、固有感覚と平面胴体速度指令のみを条件とする軽量事前ポリシーに蒸留。 - 事前ポリシーをマルチタスクRLで微調整。 - 目標条件付きタスク:任意のコマンドを追従。 - 参照誘導タスク:人間データを追従し、明示的なスタイル正則化として機能。 - これにより、操縦コマンドから協調的な全身動作を再構築。

4. どうやって有効だと検証した?

- 3体の人型ロボット(Boston Dynamics Atlas R1, Atlas D1, Unitree G1)で実験。 - 屋内・屋外環境でのユーザー操作による直接的な歩行を含む実世界シナリオでロバストな性能を実証。 - 階層制御スタックの歩行層として統合。 - 人間データなしで学習したTabula Rasa RLポリシーとのベンチマークおよびアブレーション研究で有効性を確認。

5. 議論はある?

- 本フレームワークは軽量で展開可能なポリシーを生成し、人間の歩行特性を保持しつつロバストで完全に操縦可能。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- Tabula Rasa RL policies(人間データなしで学習したRLポリシー) - 人間動作データを用いた先行研究(具体的な論文名は要旨からは不明) - 人型ロボットの歩行制御に関する一般的なRL手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mike Zhang, Dongho Kang, Kevin Bergamin, Nicola Burger, Robin Deits, Jonathan Foster, Bilal Hammoud, Katie Hughes, Francesco Iacobelli, Twan Koolen, M. Eva Mungai, Zach Nobles, Shane Rozen-Levy, Jean Pierre Sleiman, Fangzhou Yu, Yunbo Zhang, Alfred Rizzi, Jessica Hodgins, Scott Kuindersma, Yeuhi Abe, Sylvain Bertrand, Farbod Farshidian

分類: cs.RO

原文アブストラクト

Despite recent advances in humanoid locomotion, controllers optimized for command tracking and robustness tend to produce mechanical gaits, whereas controllers tied to human motion data often fail to generalize to commands outside the data distribution. This work introduces a learning framework that balances these competing objectives to synthesize real-time steerable, robust, and biomimetic locomotion policies from human data. Using an in-house curated locomotion dataset covering diverse speeds and directions, we first learn a natural locomotion prior policy through a teacher-student distillation process. Specifically, we train a full-body reference-conditioned policy with Reinforcement Learning (RL), then distill it into a lightweight prior policy conditioned solely on proprioception and a planar torso-velocity steering command. Next, we fine-tune the prior policy with multi-task RL to expand command coverage and robustness beyond the data distribution, pairing a goal-conditioned task that tracks arbitrary commands with a reference-guided task that tracks the human data as an explicit style regularizer. We validate our framework on three humanoid robots: the Boston Dynamics Atlas R1, Atlas D1, and Unitree G1. Experimental results demonstrate robust performance across real-world scenarios, including direct user-controlled locomotion in indoor and outdoor environments, and integration as the locomotion layer within hierarchical control stacks. Benchmarks against Tabula Rasa RL policies trained without human data and ablation studies confirm that our framework yields a lightweight, deployable policy that reconstructs coordinated whole-body behavior from a steering command, retaining the human gait characteristics while remaining robust and fully steerable.

関連論文

PR本紙発行元 EmplifAI