世界モデルを参照軌道として用いる高速運動適応
World Models as Reference Trajectories for Rapid Motor Adaptation
世界モデルの予測を暗黙の参照軌道として使う二重制御フレームワークを提案し、強化学習による長期報酬最大化と高速な潜在制御による運動実行を分離することで、動特性変化に素早く適応する手法を実現した。
著者: Carlos Stein Brito, Daniel McNamee
分類: cs.LG, cs.AI, cs.RO, cs.SY, eess.SY
原文アブストラクト
Deploying learned control policies in real-world environments poses a fundamental challenge. When system dynamics change unexpectedly, performance degrades until models are retrained on new data. We introduce Reflexive World Models (RWM), a dual control framework that uses world model predictions as implicit reference trajectories for rapid adaptation. Our method separates the control problem into long-term reward maximization through reinforcement learning and robust motor execution through rapid latent control. This dual architecture achieves significantly faster adaptation with low online computational cost compared to model-based RL baselines, while maintaining near-optimal performance. The approach combines the benefits of flexible policy learning through reinforcement learning with rapid error correction capabilities, providing a principled approach to maintaining performance in high-dimensional continuous control tasks under varying dynamics.
関連論文
- チャンク型VLAマニピュレーションポリシーの学習と実機展開のためのSim-to-Real統合パイプラインsim2real
- 運動学を超えて:筋駆動模倣学習のためのシミュレーション忠実度ベンチマークsim2real
- CRISP: 多様な形状と接触ソルバを備えた接触リッチロボットシミュレーション基盤sim2real
- 同じ世界、異なる知識:孤立評価が世界モデルの修復を誤判定するときsim2real
- DEXTERA: 単一画像から実機展開可能な巧みなマニピュレーションへ向けたReal-to-Sim-to-Realsim2real
- 単一スキャンからのガウシアンスプラッティングによる実演合成と視覚運動ポリシー学習sim2real