日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
制御arXiv:2607.26370v1

未知ダイナミクス追跡のための自己適応学習とモデル予測制御:無悔後悔を実現

Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

シェア:XThreadsFacebookLINEはてブBluesky

未知の目標ダイナミクスを追跡するための自己適応オンライン学習とモデル予測制御を提案し、切り替わる目標挙動に対して複数の予測器を適応的に選択することで、期待値で有限時間の準最適性を保証する。

著者: Atharva Navsalkar, Hongyu Zhou, Vasileios Tzoumas

分類: cs.RO, cs.LG, eess.SY

原文アブストラクト

We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to track, pursue, or avoid collision with moving landmarks, objects, humans, etc., whose dynamics are unknown. Our method simultaneously learns multiple predictors from scratch, via self-supervised, one-shot, and computationally efficient learning, and adaptively selects the best one to match the observed target behavior. The method enjoys finite-time near-optimality guarantees in expectation, characterized as a function of the learning error of the target dynamics and the frequency that the target dynamics switch. In the absence of both error and switching, the method asymptotically matches the optimal non-causal control policy that knows a priori the target dynamics, i.e., the method enjoys no regret in expectation. In the presence of learning errors and switching, the method degrades gracefully, \eg when there are errors and no switching, the average regret is proportional to the average learning error and switching times. To prove these guarantees, a novel technical approach is required compared to the existing works that employ RFF-based online learning. We validate our method in Crazyflie simulations and hardware experiments, across target trajectories that vary from structured to random to adversarial, in comparison to non-stochastic, kernel-based, and neural-network-based methods for online learning.

関連論文