日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
継続学習/全身制御arXiv:2610.04231

ヒューマノイド全身運動の継続学習

Continual Humanoid Motion Learning

シェア:XThreadsFacebookLINEはてブBluesky

過去データを再学習せずに新しいスキルを順次獲得できるヒューマノイド全身制御器を提案し、類似度に基づくLoRA-PNNで破滅的忘却を防ぎつつ実機Unitree G1に展開した。

詳しい要約

1. どんなもの?

- ヒューマノイドの全身運動制御を対象に、逐次タスクストリームから過去データを再訪せずにスキルを獲得する continual learning を研究。 - 単一のコントローラが壊滅的忘却を起こさずに新規スキルを学ぶことを目指す。 - Similarity-guided LoRA-PNN を提案。progressive neural network (PNN) をベースに、低ランク適応 (LoRA) で知識を再利用。 - 二段階の運動類似度 (dynamic time warping と optimal transport を組み合わせ) で、どの過去スキルを土台にし、どれだけ新規容量を割り当てるか決定。 - 6つの逐次スキルカテゴリで評価し、物理ロボット Unitree G1 に展開。

2. 先行研究と比べてどこがすごい?

- 従来のヒューマノイド全身コントローラはオフラインで訓練後に固定され、新スキルを教えると既存スキルが劣化する問題があった。 - 提案手法は PNN により構造的に壊滅的忘却を防止しつつ、LoRA で軽量に知識を再利用。 - 類似度に基づく容量割り当てにより、forward transfer が 0.125 対 0.079 と向上。 - 全手法中で最高の平均精度を達成し、訓練可能パラメータを最大 94.5%、訓練時間を 40.8% 削減。 - sim-to-sim 転移 96.13% を達成し、物理 Unitree G1 に展開可能。

3. 技術・手法の肝は?

- progressive neural network (PNN) をベースに、タスクごとに新規サブネットワークを追加し、過去の重みを凍結することで忘却を防止。 - 新規スキル獲得時には LoRA (low-rank adaptation) を用いて既存知識を軽量に再利用。 - 二段階の運動類似度尺度を導入: dynamic time warping で時系列間の類似度を計算し、optimal transport で集約。 - この類似度に基づき、どの過去スキルを土台にするか、どれだけ新規容量を割り当てるかを決定。 - これにより forward transfer と効率を両立。

4. どうやって有効だと検証した?

- 6つの逐次スキルカテゴリで実験を実施。 - forward transfer が 0.125 (比較手法 0.079) と最高。 - 全手法中で最高の平均精度を達成。 - 訓練可能パラメータを最大 94.5%、訓練時間を 40.8% 削減。 - sim-to-sim 転移 96.13% を達成し、物理ロボット Unitree G1 に展開。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- progressive neural network (PNN) - LoRA (low-rank adaptation) - dynamic time warping - optimal transport - continual learning - sim-to-sim transfer - Unitree G1

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhewen He, Hao Huang, Geeta Chandra Raju Bethala, Chong Yu, Tao Chen, Anthony Tzes, Yi Fang

分類: cs.RO

原文アブストラクト

Humanoid whole-body controllers can now track a diverse set of dynamic motions, but they are typically trained offline and then frozen, so teaching such a controller a new skill tends to erode the skills it already mastered. We study continual learning for humanoid whole-body motion, where a single controller must acquire skills from a sequential task stream without revisiting past data. We introduce Similarity-guided LoRA-PNN, a progressive neural network (PNN) policy that prevents catastrophic forgetting by construction while reusing knowledge across skills through lightweight low-rank adaptation. A two-level motion-similarity measure, built from dynamic time warping aggregated by optimal transport, decides which prior skill to build on and how much new capacity to allocate, yielding strong forward transfer and large efficiency gains. Across six sequentially learned skill categories, our similarity-guided LoRA policy attains the best forward transfer (0.125 vs. 0.079) and the highest average accuracy among all methods, while saving up to 94.5% of trainable parameters and 40.8% of training time. The resulting controller reaches 96.13% sim-to-sim transfer and is deployed on a physical Unitree G1. Our code is available at https://anonymous.4open.science/r/continual-humanoid-learning-35D3.

PR本紙発行元 EmplifAI