日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.01260

PROMO: 四足歩行ロボットのための選好条件付き多目的強化学習

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

シェア:XThreadsFacebookLINEはてブBluesky

単一の方策に実行時の選好を入力することで、指令追従・安定性・省エネルギーのトレードオフを調整可能にし、シミュレーションと実機Unitree Go2で有効性を示した。

詳しい要約

1. どんなもの?

- 四足歩行ロボットのための強化学習手法「PROMO」を提案。 - 歩行にはcommand tracking、stability、energy efficiencyなど相反する目的が存在。 - 従来のRLは訓練時に固定のスカラー報酬へ優先度を埋め込む。 - PROMOは選好(preference)を実行時の明示的入力とし、単一のlocomotion policyを条件付ける。 - 意味論的multi-objective approachで、operator intentと報酬整形項を分離。

2. 先行研究と比べてどこがすごい?

- 固定目的のcontroller、multi-objective baseline、個別訓練のspecialistと比較。 - 単一のdeployable policyでobjective specializationとrobustnessを達成。 - 従来のmulti-objective RLはofflineでのPareto-set構築が主だが、実行時interfaceへ拡張。 - 選好変更のみでenergy、position error、attitude deviationを改善可能。

3. 技術・手法の肝は?

- preference-conditioned multi-objective RL。 - policyをdeployment facing preferencesで条件付け。 - embodiment-specific locomotion priorsは固定し、operator intentと報酬整形項を分離。 - 単一policyで多様な選好に対応。 - 詳細なアルゴリズムは要旨からは不明。

4. どうやって有効だと検証した?

- シミュレーションで100個の選好をサンプリング。 - 67 behaviorsがexact Pareto dominanceでnon-dominated。 - 平均preference-objective correlationは0.843。 - 広いPareto coverageと予測可能な選好応答を示す。 - 同一policyをUnitree Go2へzero-shot転移。 - 選好変更のみでspecific energy最大30.4%減、position error 38.7%減、peak body-attitude deviation 59.0%減(balanced preference比)。

5. 議論はある?

- preference-conditioned multi-objective RLが適応的legged locomotionの実用的runtime interfaceとなることを示す。 - offline Pareto-set構築を超える役割を拡張。 - 限界や課題は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: fixed-objective controllers、multi-objective baselines、independently trained specialists。 - 関連手法: multi-objective reinforcement learning、Pareto dominance、preference-conditioned RL。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger

分類: cs.RO, cs.AI, cs.HC, cs.LG, eess.SY

原文アブストラクト

Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.

関連論文

PR本紙発行元 EmplifAI