QPベースのエンドツーエンド方策:ロバスト制御とロボット学習の統一的な視点
End-to-end QP-based policies: A unified perspective on robust control and robot learning
二次計画法(QP)をベースにしたエンドツーエンドの方策フレームワークを提案し、モデルベース制御の透明性を保ちつつ、ドメインランダム化による自動調整やブラックボックス方策構築を可能にした。シミュレーションと実機でロバスト性や未知ダイナミクス下の制御を検証した。
著者: Fausto Vega, Priyanka Supraja Balaji, Chase Dunaway, Joe Koszut, Jon Arrizabalaga, Zachary Manchester
分類: cs.RO
原文アブストラクト
We present an end-to-end QP-based policy framework that enables systematic policy construction with minimal domain-specific design, while preserving the transparency and interpretability of model-based control. The proposed policy representation supports both domain-randomized model-based auto-tuning, where policy parameters are optimized over distributions of disturbances and model variations, and black-box policy construction, where the problem is formulated in terms of a (possibly) unknown model without requiring explicit notions of states, inputs, or the underlying system dynamics. We establish connections to existing policy representations and control paradigms, including robust control, multilayer perceptrons, and robot learning, and interpret our end-to-end QP policies as a common abstraction of these approaches. We validate the resulting framework in both simulation and hardware, demonstrating a broad range of applications spanning robustness, automatic policy tuning, and control under unknown system dynamics.