日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.10297

階層強化学習による多様な地形での四脚歩行のエネルギー効率の良い歩容適応

Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains

シェア:XThreadsFacebookLINEはてブBluesky

高周波の関節運動制御と低周波の歩容適応を分離した階層強化学習により、速度に応じた歩容切替とエネルギー効率を両立し、実機四脚ロボットへゼロショット転移した研究。

詳しい要約

1. どんなもの?

- 四足歩行ロボットのエネルギー効率を高めるための階層型強化学習(HRL)フレームワークを提案。 - 高頻度ポリシーが関節レベルの運動実行を担当し、低頻度ポリシーが歩容適応とCoT最小化を担当。 - 3段階のIsaacベース訓練により、ゼロショットのsim-to-real転移を実現。 - 速度に応じた自動歩容適応(低速でpacing、高速でtrotting)を示す。 - シミュレーションで単一ポリシーや階層型ベースラインと比較し、広い速度範囲でCoTを低減。 - 平坦、不整地、傾斜地でのロバストな歩行を維持し、Unitree AlienGoで実機展開を実証。

2. 先行研究と比べてどこがすごい?

- 従来のend-to-end RLポリシーでは、歩容生成、運動実行、エネルギー最適化が密結合し、報酬設計に敏感でエネルギー効率とロバスト性の両立が困難。 - 提案するHRLはこれらを分離し、低頻度ポリシーで明示的にCoTを最小化。 - 単一ポリシーや代表的な階層型歩行ベースラインと比較して、広い指令速度範囲でCoTを低減。 - ゼロショットsim-to-real転移を実現し、追従精度、ロバスト性、エネルギー効率を向上。

3. 技術・手法の肝は?

- 階層型強化学習(HRL)フレームワークを採用。 - 高頻度ポリシー:安定かつロバストな関節レベル運動実行を担当。 - 低頻度ポリシー:歩容適応を担当し、コスト・オブ・トランスポート(CoT)を明示的に最小化。 - 3段階のIsaacベース訓練手順を実施。 - 速度依存の自動歩容適応を学習(低速でpacing、高速でtrotting)。

4. どうやって有効だと検証した?

- シミュレーションで代表的な単一ポリシーおよび階層型歩行ベースラインと比較。 - 広い指令速度範囲でCoTの低減を実証。 - 平坦、不整地、傾斜地でのロバストな歩行を確認。 - 物理的なUnitree AlienGo四足ロボットへのゼロショット展開により実用性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:単一ポリシーおよび階層型歩行ベースライン(具体的名称は要旨に記載なし)。 - 関連手法:end-to-end RLポリシー、階層型強化学習(HRL)。 - 同分野の定番:四足歩行ロボットの歩容制御、sim-to-real転移、エネルギー効率最適化に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ammar Issa, Anubhav Singh, Anton Tsaritsin, Sergey Kolyubin

分類: cs.RO, cs.LG

原文アブストラクト

While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.

関連論文

PR本紙発行元 EmplifAI