日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
外骨格/強化学習/sim2realarXiv:2609.28027

Sim-to-real強化学習による速度適応型股関節外骨格制御ポリシーの学習

Learning a Speed-adaptive Hip Exoskeleton Control Policy Via Sim-to-real Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションで歩行速度に応じたアシストタイミングを強化学習し、実機の股関節外骨格に転移させ、オンライン選好学習でアシスト強度を個人適応させるフレームワークを提案した。

詳しい要約

1. どんなもの?

本論文は、歩行速度に適応する股関節用 exoskeleton の制御ポリシーを、sim-to-real reinforcement learning (RL) とオンライン preference learning を統合して学習するフレームワークを提案する。assistance timing は simulation 上で human musculoskeletal models を用いた RL により学習し、そのポリシーを distill して実機の hip exoskeleton に搭載する。assistance magnitude は Gaussian-process-based preference learning により、ユーザーの pairwise 比較から個人化する。

2. 先行研究と比べてどこがすごい?

既存の online optimization 法は sample-inefficient で、assistive torque profile 全体を最適化するために多数の human-in-the-loop (HIL) 評価を要する。sim-to-real RL は有望だが個人の嗜好を直接扱えない。本研究は timing の simulation 学習と magnitude の実世界 preference learning を分離し、オンライン最適化空間を大幅に削減する点が先行研究と異なる。

3. 技術・手法の肝は?

assistance timing を simulation で学習するため、human musculoskeletal models と varying walking speeds を用いて RL ポリシーを訓練する。学習済みポリシーを distill し、onboard sensory observations を用いて physical hip exoskeleton に実装する。assistance magnitude は Gaussian-process-based preference learning により、ユーザーの pairwise 比較を通じて個人化する。これにより timing 学習と magnitude 最適化を分離する。

4. どうやって有効だと検証した?

human-subject experiments を実施し、varying walking speeds にわたり、より少ない real-world evaluations で個人化された assistive torque profiles を効率的に同定できることを示した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として sim-to-real reinforcement learning、human-in-the-loop (HIL) optimization、Gaussian-process-based preference learning、human musculoskeletal models を用いた exoskeleton control の文献を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bin Li, Zhimin Hou, Jiacheng Hou, Zenian Liang, Tong Wu, Teng Ma, Chenglong Fu

分類: cs.RO

原文アブストラクト

Providing personalized exoskeleton assistance across varying walking speeds remains challenging. Existing online optimization methods are sample-inefficient, requiring extensive human-in-the-loop (HIL) evaluations to optimize the entire assistive torque profile. Sim-to-real reinforcement learning (RL) offers a promising alternative but cannot directly account for individual user preferences. We propose a framework integrating sim-to-real RL with online preference learning for personalized exoskeleton assistance. Specifically, assistance timing is learned in simulation by training RL policies with human musculoskeletal models across varying walking speeds. The learned policies are then distilled and deployed on a physical hip exoskeleton using onboard sensory observations. Gaussian-process-based preference learning further personalizes the assistance magnitude through pairwise user comparisons. By decoupling assistance timing learning in simulation from magnitude optimization in real-world experiments, our framework substantially reduces the online optimization space. Human-subject experiments demonstrate efficient identification of personalized assistive torque profiles across varying walking speeds with fewer real-world evaluations.

PR本紙発行元 EmplifAI