Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the real-time tracking control problem under unknown target dynamics that may switch in structured, stochastic, or adversarial manners. The authors propose a novel approach integrating adaptive online learning with model predictive control, which concurrently learns multiple dynamic predictors and adaptively selects the best-performing model based on observations. The method employs a self-supervised, single-shot, computationally efficient multi-model learning mechanism that avoids reliance on conventional techniques such as random Fourier features. Theoretical analysis demonstrates that the resulting no-regret controller asymptotically achieves non-causal optimal performance in the absence of disturbances and guarantees finite-time near-optimality; under disturbances, it ensures gracefully degraded performance. Extensive simulations and hardware experiments on the Crazyflie platform confirm its consistent and significant superiority over non-stochastic, kernel-based, and neural network baselines across diverse trajectories.
📝 Abstract
We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to track, pursue, or avoid collision with moving landmarks, objects, humans, etc., whose dynamics are unknown. Our method simultaneously learns multiple predictors from scratch, via self-supervised, one-shot, and computationally efficient learning, and adaptively selects the best one to match the observed target behavior. The method enjoys finite-time near-optimality guarantees in expectation, characterized as a function of the learning error of the target dynamics and the frequency that the target dynamics switch. In the absence of both error and switching, the method asymptotically matches the optimal non-causal control policy that knows a priori the target dynamics, i.e., the method enjoys no regret in expectation. In the presence of learning errors and switching, the method degrades gracefully, \eg when there are errors and no switching, the average regret is proportional to the average learning error and switching times. To prove these guarantees, a novel technical approach is required compared to the existing works that employ RFF-based online learning. We validate our method in Crazyflie simulations and hardware experiments, across target trajectories that vary from structured to random to adversarial, in comparison to non-stochastic, kernel-based, and neural-network-based methods for online learning.
Problem

Research questions and friction points this paper is trying to address.

unknown dynamics
target tracking
switching behavior
no regret
online learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-adaptive learning
model predictive control
no-regret tracking
online learning
switching dynamics
🔎 Similar Papers
2024-10-04arXiv.orgCitations: 0