Interleaved Projected Gradient Descent for Safe Imitation Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses safe imitation learning under state and input constraints, where existing methods struggle to simultaneously ensure safety and control performance. To this end, it proposes an interleaved projected gradient descent approach that alternates between standard gradient updates and safety set projections during training. At inference, the method relies solely on pure network forward propagation without requiring additional safety filters, supported by theoretical guarantees for asymptotic constraint satisfaction and loss bounds. Evaluated in a nonlinear autonomous racing simulation, the proposed approach reduces the constraint violation rate from 15% to 1% while achieving lap times comparable to unconstrained baselines and demonstrating robustness to hyperparameter variations. Ultimately, this work realizes a synergistic optimization of theoretical safety guarantees and practical control performance.
📝 Abstract
We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of $k$ safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting $k$ grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from $15\%$ to $1\%$, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
Problem

Research questions and friction points this paper is trying to address.

Imitation Learning
Safe Control
State and Input Constraints
Neural Network Policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imitation Learning
Projected Gradient Descent
Safety Constraints
Neural Network Control
Autonomous Racing
💼 Related Jobs
No related jobs found.