日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.28878

閉ループシステムモデリングによるオンラインSim-to-Real適応

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

シェア:XThreadsFacebookLINEはてブBluesky

既存の制御器への参照コマンドをオンラインで適応させることで、シミュレーションと実機のギャップによる追従誤差を低減するフレームワークOSRAMを提案。

詳しい要約

1. どんなもの?

- 本論文は、sim-to-real転送における残差動力学ミスマッチによる追従精度低下を解決するOSRAMを提案する。 - OSRAMは、既存のコントローラへの参照コマンドを適応させるフレームワークである。 - 展開されたロボットとそのポリシーを統合された閉ループ動力学システムとして扱い、追従観測からタスクレベルのコマンド応答挙動を直接学習する。 - 閉ループ動力学モデルは、シミュレーションでランダム化された動力学にわたってメタ学習され、展開後に限られた実世界インタラクションで迅速にファインチューニングされる。 - 適応されたモデルは、基礎となる制御ポリシーを変更せずに、将来の参照コマンドを最適化するために使用される。

2. 先行研究と比べてどこがすごい?

- 従来のsim-to-real転送では、残差動力学ミスマッチを修正するために、システム動力学の同定、制御ポリシーの適応、またはシミュレーションへの追加訓練とファインチューニングが必要であり、これらは大量のデータと計算を要する。 - OSRAMは、代わりに既存のコントローラへの参照コマンドを適応させることで、ポリシーのファインチューニングや完全な物理動力学の同定を回避する。 - これにより、データと計算の負担を軽減しつつ、残差sim-to-real追従誤差を低減する実用的な代替手段を提供する。

3. 技術・手法の肝は?

- OSRAMは、展開されたロボットとそのポリシーを統合された閉ループ動力学システムとして扱う。 - 追従観測からタスクレベルのコマンド応答挙動を直接学習する。 - 閉ループ動力学モデルは、シミュレーションでランダム化された動力学にわたってメタ学習される。 - 展開後、限られた実世界インタラクションを用いてモデルを迅速にファインチューニングする。 - 適応されたモデルを用いて将来の参照コマンドを最適化するが、基礎となる制御ポリシーは変更しない。

4. どうやって有効だと検証した?

- シミュレーションとハードウェア上で、bipedal velocity trackingとloco-manipulationを対象にOSRAMを評価した。 - 結果は、閉ループモデリングが未知の動力学下で予測と追従精度を改善することを示した。 - オンライン参照適応が、異なる制御目的とハードウェア構成にわたって残差sim-to-real追従誤差を低減することを示した。 - これらの結果は、ロボット-ポリシー閉ループの挙動を適応させることが、sim-to-real転送のためのポリシーのファインチューニングや完全な物理動力学の同定に対する実用的な代替手段を提供することを実証している。

5. 議論はある?

- 要旨からは、OSRAMの限界や潜在的な問題点についての具体的な議論は不明である。 - ただし、結果は、閉ループモデリングとオンライン参照適応が、異なる制御目的とハードウェア構成にわたって有効であることを示している。 - また、ポリシーのファインチューニングや完全な物理動力学の同定の代替として実用的であると主張している。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、sim-to-real転送、domain randomization、system identification、meta-learning、model predictive control (MPC)、reinforcement learning (RL) などが挙げられる。 - 具体的な論文名は要旨からは不明である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuhao Huang, Samuel A. Moore, Boyuan Chen

分類: cs.RO

原文アブストラクト

Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.

関連論文

PR本紙発行元 EmplifAI