trust-region optimization

Designs and analyzes optimization algorithms and update rules that constrain each parameter or distribution update to lie inside a trust region, including methods to adapt the trust-region radius and to penalize deviations from a prior. Implements Lyapunov-constrained trust-region conditions to guarantee per-iterate feasibility or decrease, and develops sequential and multi-agent trust-region strategies and practical trust-region update mechanisms.

trust-regionoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

Aug 14, 2025
CL
Chak Lam Shek
🏛️ University of Maryland | University of Southern California

To address uncoordinated policy updates and training instability in heterogeneous multi-agent reinforcement learning (MARL), this paper proposes a joint constrained optimization framework that dynamically allocates per-agent KL-divergence trust-region thresholds. Unlike HATRPO, which imposes a uniform KL constraint across all agents, we introduce two adaptive threshold allocation mechanisms: (1) HATRPO-W, an analytically derived solution grounded in the Karush–Kuhn–Tucker (KKT) optimality conditions; and (2) HATRPO-G, a greedy scheduling algorithm guided by an improvement-to-divergence ratio criterion. Both methods enforce coordinated policy updates under a global KL budget. Empirical results demonstrate that HATRPO-W and HATRPO-G outperform baseline HATRPO by over 22.5% in average task performance. Notably, HATRPO-W achieves faster convergence and significantly enhanced training stability. These findings underscore the critical role of dynamic, agent-specific trust-region adaptation in improving both efficiency and robustness of heterogeneous MARL training.

Addressing heterogeneous-agent trust region threshold allocationImproving MARL convergence and reward via flexible KL divergence schedulingOptimizing multi-agent policy updates with joint constraints

This work addresses the challenges of high-dimensional black-box constrained optimization, where function evaluations are expensive, gradient information is unavailable, and the feasible region is complex. The authors propose a Bayesian optimization method that integrates a penalty function with a trust-region mechanism. By incorporating constraints into an unconstrained formulation via penalty terms, the approach constructs a local Gaussian process surrogate model and performs sampling within a dynamically adjusted trust region using the expected improvement criterion, thereby effectively balancing exploration and exploitation. Experimental results demonstrate that the proposed method consistently achieves high-quality feasible solutions with significantly fewer function evaluations than state-of-the-art approaches across multiple high-dimensional synthetic and real-world constrained optimization problems, yielding notable improvements in both sample efficiency and optimization stability.

black-boxconstrained optimizationexpensive evaluations

Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference

Aug 17, 2025
DB
Denis Blessing
🏛️ Karlsruhe Institute of Technology | NVIDIA | Zuse Institute Berlin | Microsoft Research New England | Cornell University

In stochastic optimal control and inference, approximating path-space measures becomes challenging when the target and prior measures differ substantially, leading to gradient mismatch and poor convergence in conventional optimization. To address this, we propose a trust-region-based iterative constrained optimization framework—the first to incorporate trust-region mechanisms into path-space measure transport—enabling geometrically annealed, progressive approximation. Explicit feasibility constraints ensure stable updates at each iteration, while a systematic annealing step-size selection rule is introduced. Our method unifies gradient-based optimization, optimal transport theory, and robust constraint principles. Experiments on diffusion sampling, transition path simulation, and diffusion model fine-tuning demonstrate significant improvements in training stability and convergence speed. The approach exhibits strong generalization across diverse tasks, validating both its effectiveness and broad applicability.

Approaching target measure via geometric annealingImproving performance in diffusion-based tasksSolving stochastic optimal control with trust regions

This work addresses the issue of exponentially high variance in joint advantage estimation under sequential policy updates in cooperative multi-agent reinforcement learning, which leads to unstable training. The paper presents the first theoretical analysis of advantage variance under sequential update schemes and introduces a novel clipped objective function that integrates trust-region methods with importance sampling to effectively bound the magnitude of advantage fluctuations. The proposed approach yields two algorithms that guarantee monotonic policy improvement and converge to an ε-Nash equilibrium at a sublinear rate. Empirical results demonstrate that the method significantly outperforms existing baselines across three standard MARL benchmarks, achieving substantially lower advantage estimation variance and more stable convergence behavior.

advantage estimationmulti-agent reinforcement learningsequential updates

This paper addresses distributed optimization in multi-agent cyber-physical systems under adversarial conditions—specifically, malicious neighbors and heterogeneous, biased local objective functions. Method: We propose the first resilience-oriented optimization framework grounded in stochastic trust relationships. The approach integrates stochastic graph modeling, robust consensus protocols, and trust-weighted gradient updates. Convergence is rigorously established via martingale analysis, guaranteeing both mean-square and almost-sure convergence to the global optimum—even when malicious agents constitute a majority—and yielding an expected convergence rate of $O(1/k)$. Contribution/Results: Our framework breaks the conventional “majority-benign” assumption in resilient distributed optimization. Experiments demonstrate stable convergence even when over 50% of agents are malicious—a regime where existing methods fail. This work establishes a verifiably resilient optimization paradigm for highly adversarial multi-robot networks, enabling provable performance guarantees under extreme attack conditions.

Multi-Robot SystemsResilient Network OptimizationTrust Issues

Latest Papers

What's happening recently
View more

This study addresses the challenging problem of stochastic nonlinear optimization with deterministic equality constraints under heavy-tailed noise characterized by unbounded variance. To this end, we propose TR-SSQP, a method grounded in a trust-region framework and stochastic sequential quadratic programming. By leveraging a normal-tangential decomposition to balance optimality and feasibility, combined with radius normalization and Polyak momentum for stable updates, the algorithm achieves robust performance. The core contribution lies in establishing the first global almost-sure convergence theory without resorting to gradient clipping, thereby filling a critical theoretical gap in this setting. Numerical experiments demonstrate that the proposed approach significantly outperforms existing constrained stochastic optimization algorithms.

constrained stochastic optimizationequality constraintsheavy-tailed noise

This work addresses the degraded convergence in distributed first-order optimization caused by gradient compression and the lack of tight theoretical comparisons among existing error feedback algorithms (EF/EF21). To resolve these issues, the authors develop a unified convergence analysis framework. By constructing optimal Lyapunov functions tailored to EF and EF21 and identifying their respective optimal step sizes, they establish the first tight convergence guarantees that are independent of the number of nodes and match the optimal convergence rate of single-agent methods. This analysis not only clarifies the distinct convergence behaviors of EF and EF21 under arbitrary network scales but also characterizes their optimal rates without relying on specific communication graph structures, thereby significantly enhancing both the clarity and practical utility of the theoretical understanding.

communication costconvergence analysisdistributed optimization

This work addresses the high computational cost of traditional interior-point trust-region methods for constrained nonconvex optimization, which stems from repeatedly solving trust-region subproblems. To overcome this limitation, we propose an approximate first-order interior-point trust-region framework that maintains an approximate projection operator via low-rank updates, thereby avoiding explicit subproblem solves and Hessian evaluations. By integrating a gradient-based negative curvature routine, the method efficiently approximates both first- and second-order Karush–Kuhn–Tucker (KKT) points using only first-order information. Theoretical analysis establishes feasibility and convergence guarantees for the iterates, while numerical experiments demonstrate up to a 2.48× speedup over state-of-the-art methods on large-scale problems, significantly enhancing scalability.

constrained nonconvex optimizationinterior-point trust-regionKarush–Kuhn–Tucker points

This work addresses the high variance in advantage estimation caused by non-stationary teammate policies in multi-agent reinforcement learning, which undermines the effectiveness of ratio-based trust-region methods such as MAPPO and MASPO. To mitigate this issue, the authors propose MARS, a novel policy optimization objective that replaces conventional additive ratio clipping or soft penalty mechanisms with a multiplicative symmetric geometric barrier within the centralized training with decentralized execution (CTDE) framework. This design imposes unbounded penalties on probability ratios approaching zero while preserving informative gradients, thereby preventing policy collapse and vanishing gradients. Empirical evaluation across 47 tasks spanning eight benchmark environments demonstrates that MARS consistently matches or outperforms existing methods, and ablation studies confirm the critical role of the symmetric geometric barrier in its performance gains.

multi-agent reinforcement learningnon-stationaritypolicy optimization

Standard PPO struggles to transfer policies effectively in non-stationary environments due to its geometrically unaware local updates and excessive regularization that overly suppresses necessary policy adjustments. This work proposes Gaussian Trust Region (GTR) optimization, which dynamically reshapes the trust region using a Gaussian kernel to enable adaptive, large-magnitude updates along high-advantage directions while preserving local stability. Additionally, GTR introduces dynamically blended Gaussian anchor policies to mitigate variance caused by outdated reference policies. The resulting framework is architecture-agnostic and demonstrates consistent and significant improvements over strong baselines across diverse domains—including gaming, robotic control, open-world exploration, and language model post-training—highlighting its generality and robustness in non-stationary settings.

behavior transitionsgeometric awarenessnon-stationary environments

Hot Scholars

JL

Javad Lavaei

Associate Professor, UC Berkeley
OptimizationMachine LearningControlEnergy
RK

Rolf Krause

Full Professor, KAUST
Numerical Solution of PDEsMachine LearningMultigrid/Domain DecompositionContact Problems
YF

Yuchen Fang

University of California, Berkeley
constrained stochastic optimizationhigh-dimensional statistics
SN

Sen Na

Assistant Professor in Industrial and Systems Engineering, Georgia Tech
statisticsoptimizationmachine learningdata science