日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.21307

安定性を考慮した残差強化学習によるマニピュレータ外乱補償フレームワーク

Stability-aware Residual Reinforcement Learning Framework for Robotic Manipulator Disturbance Compensation

シェア:XThreadsFacebookLINEはてブBluesky

外乱オブザーバと強化学習を組み合わせ、残差を学習して補償する手法を提案。ISS解析に基づく行動制約で安定性を保証し、6自由度マニピュレータで追従誤差を最大38%削減した。

詳しい要約

1. どんなもの?

- ロボットマニピュレータの外乱補償のためのフレームワーク。 - 解析的な外乱オブザーバ(DOB)と強化学習(RL)ポリシーを組み合わせた残差強化学習DOB。 - 決定論的ベースラインが信頼できる領域で動作し、RLポリシーがモデルで捉えられない残差を明示的に補償。 - 外乱認識を可能にする推定器ネットワークを導入し、観測履歴を特権的な外乱コンテキストと整合。 - 入力-状態安定性(ISS)解析に基づく状態依存の行動境界をRLポリシーに課し、閉ループの追従誤差を認定エンベロープに閉じ込める。

2. 先行研究と比べてどこがすごい?

- 従来のコントローラやDOBはパラメータ不確かさ、非線形摩擦、複合外乱に悩まされる。 - 提案手法は解析的オブザーバとRLを組み合わせ、残差をRLで補償することでこれらの問題に対処。 - 外乱レジームごとに潜在空間を組織化し、外乱遷移に迅速に適応可能。 - ISS解析に基づく行動境界により、任意のポリシー出力に対して安定性を保証。 - ゼロショットsim-to-real転移や未学習の外乱下でも追従誤差を大幅に削減。

3. 技術・手法の肝は?

- 解析的DOBとRLポリシーを併用する残差強化学習フレームワーク。 - 推定器ネットワークが観測履歴を特権的な外乱コンテキストと整合させ、外乱レジームごとに潜在空間を組織化。 - ISS解析から状態依存の行動境界を導出し、RLポリシーに強制。 - これにより閉ループの追従誤差が認定エンベロープ内に収まることを証明。 - ベースラインは信頼領域で動作し、RLは残差のみを対象とする。

4. どうやって有効だと検証した?

- 6自由度マニピュレータでの実験。 - 外乱推定と追従性能の一貫した改善を確認。 - 実ハードウェアでのゼロショットsim-to-real転移で追従誤差27.8%削減。 - 訓練時に観測されなかったベース振動外乱下で38.0%削減。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、Disturbance Observer (DOB)、Residual Reinforcement Learning、Input-to-State Stability (ISS) が挙げられる。 - 同分野の定番として、Model Predictive Control (MPC) や Domain Randomization も参考になる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jihong Kim, Joonhyuk Kwon, Hwa Soo Kim, TaeWon Seo, Hyung-Tae Seo

分類: cs.RO

原文アブストラクト

Although conventional controllers and disturbance observers (DOBs) are the standard for precision tracking in manipulators, they suffer from parameter uncertainty, nonlinear friction, and compound disturbances. This study proposes a residual reinforcement learning DOB framework that pairs an analytical observer with an RL policy. The deterministic baseline operates within a reliable region, whereas the RL policy explicitly targets the residuals that the model cannot capture. To make this compensation disturbance-aware, an estimator network aligns the observation history with a privileged disturbance context, organizing the latent space by disturbance regime and enabling rapid adaptation across disturbance transitions. To guarantee stability, we derived and enforced a state-dependent action bound on the RL policy from an input-to-state stability (ISS) analysis such that the closed loop provably confines the tracking error to a certified envelope for arbitrary policy outputs. Experiments on a 6-DOF manipulator demonstrated consistent improvements in disturbance estimation and tracking, including a 27.8% tracking-error reduction on real hardware under zero-shot sim-to-real transfer and a 38.0% reduction under a base-vibration disturbance that was not observed during training.

関連論文

PR本紙発行元 EmplifAI