Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of learning unknown constraints, which typically relies on known dynamics or entails high-risk online exploration. To this end, we propose the CF-KKT framework, which integrates differentiable dynamics models with locally optimal demonstrations to combine the data efficiency of CIOC with the flexibility of ICRL. By leveraging Karushโ€“Kuhnโ€“Tucker (KKT) condition optimization and generating synthetic infeasible data through counterfactual actions, the proposed method recovers unknown constraints without requiring hazardous online exploration. Experimental evaluations on high-dimensional robotic control tasks demonstrate that the CF-KKT framework significantly improves both safety and data efficiency compared to offline ICRL baselines.
๐Ÿ“ Abstract
Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.
Problem

Research questions and friction points this paper is trying to address.

Constraint Learning
Learning from Demonstrations
Inverse Constrained Reinforcement Learning
Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual KKT
Constraint Learning
Differentiable Dynamics
Learning from Demonstrations
Safe Exploration
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.