日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.05855

劣駆動二足歩行ロボットの衝突回避歩行のための階層型強化学習

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

シェア:XThreadsFacebookLINEはてブBluesky

劣駆動二足歩行ロボットにおいて、高レベル方策が速度指令を出し低レベル方策が関節を制御する階層型強化学習を提案し、障害物回避歩行を実現した。

詳しい要約

1. どんなもの?

- 劣駆動2足歩行ロボットの衝突回避移動を実現するHierarchical Reinforcement Learning (HRL) フレームワーク。 - High-Level (HL) ポリシーが姿勢、36本のraycast近接測定、移動障害物状態、receding-horizon局所ゴールを観測し、10制御ステップごとに体速度指令(vx, vy, ωyaw)を出力。 - Low-Level (LL) ポリシーが速度条件付きでPD制御関節目標を追従。 - 両ポリシーをSoft Actor-Critic (SAC) で同時訓練。

2. 先行研究と比べてどこがすごい?

- 従来のプランナー併用型(SAC+A*, SAC+RRT*, SAC+APF)と比較して、静的環境で98.0%対最大78.0%、動的環境で88.0%対最大68.0%の成功率。 - 経路長はA*参照の4%以内。 - 各観測チャネルと報酬項の寄与をアブレーションで確認。

3. 技術・手法の肝は?

- HLポリシーが低次元の体速度指令を生成し、LLポリシーがそれを関節目標に変換する階層構造。 - 両ポリシーをSACでjoint training。 - 収束した歩容はタスク非依存のため凍結し、同じコマンドインタフェースで古典的プランナー(A*, RRT*, APF)により駆動可能。

4. どうやって有効だと検証した?

- ランダム化PyBullet環境で各手法100試行の評価。 - 提案手法は静的98.0%、動的88.0%の成功率。 - プランナー併用型は静的78.0%以下、動的68.0%以下。 - 経路長はA*参照の4%以内。 - アブレーションで各観測チャネルと報酬項の寄与を確認。

5. 議論はある?

- 各観測チャネルと報酬項が性能に実質的に寄与することをアブレーションで確認。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: SAC+A*, SAC+RRT*, SAC+APF。 - 関連手法: Soft Actor-Critic (SAC), Hierarchical Reinforcement Learning (HRL), A*, RRT*, APF。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., Samiran Datta, Abhay Dwivedi, Amit Shukla

分類: cs.RO, cs.AI, cs.HC, eess.SY

原文アブストラクト

A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, ω_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.

関連論文

PR本紙発行元 EmplifAI