ステップ間予測に基づくモデル情報活用型安全強化学習による二足歩行
Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
ALIPテンプレートに基づく安全証明を訓練時の報酬整形と実行時の行動フィルタに用い、人型ロボットの二足歩行の安全性を高める手法を提案した。
著者: Victor Paredes, Ayonga Hereid
分類: cs.RO
原文アブストラクト
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.