priority-constrained descent

Designs, implements, or analyzes optimization update rules that enforce a strict priority ordering among multiple objectives by computing search directions that preserve descent for higher‑priority (primary) objectives while enabling improvement on lower‑priority (secondary) objectives, typically via projection or minimal distortion of gradients. These methods include a tunable distortion strength (e.g., parameter τ) and provide mechanisms or guarantees to control how much the primary descent is altered and to ensure progress on secondary objectives without violating higher‑priority goals.

priority-constraineddescent

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the common oversight in multi-objective deep learning of ignoring the hierarchical importance among objectives, which often compromises the prioritization of primary goals while satisfying secondary constraints. To resolve this, the paper proposes Priority-Constrained Descent (PCD), the first method to explicitly model objective hierarchies by ensuring continuous descent of the primary objective while minimally perturbing the optimization trajectory to meet progress requirements of secondary objectives. A single, interpretable parameter τ governs the trade-off strength, and the approach yields scale-invariant closed-form solutions for two- to three-objective problems. Experiments on model compression and sparsification demonstrate that PCD achieves Pareto superiority, outperforming existing methods on both primary and secondary objectives, with τ effectively modulating optimization behavior.

gradient-based optimizationmulti-objective optimizationobjective hierarchy

Robot path planning and trajectory optimization are commonly formulated as optimal control problems (OCPs), yet designing appropriate trade-offs among multi-objective cost components remains challenging, and resulting solutions often lack interpretability—leading to inefficient debugging. Method: We propose the first direction-corrected cost consistency analysis framework, integrating sensitivity analysis, gradient direction projection, and expert-feedback-driven iterative reweighting optimization. Contribution/Results: This approach enables interpretable diagnostic analysis of cost components and automated weight tuning, shifting from conventional trial-and-error to goal-directed correction. It significantly improves solution rationality and task success rates while supporting adaptive objective function reconstruction with low cost and minimal samples.

Automatically tuning OCP parameters for desired correctionsBalancing multiple objective components in optimal control problemsUnderstanding impact of cost components on undesired solutions

Existing evaluation methods assess optimizer effectiveness only indirectly through the final performance of the target agent, making it difficult to evaluate the quality of individual optimization decisions. This work proposes a prioritization-based approach that requires optimizers to rank components—such as tools—according to their potential performance gain, enabling direct, step-level evaluation without costly replay or manual inspection. Leveraging the Shor dataset, which comprises 182 human-validated cross-domain scenarios, we introduce the first low-cost, scalable benchmark task tailored for harness optimizers. Experimental results demonstrate a strong correlation between an optimizer’s performance on this ranking task and its ultimate effectiveness in real multi-step optimization, thereby validating the proposed method as a reliable predictive metric.

agent performancedirect evaluationharness optimization

This work addresses the challenge of integrating augmented Lagrangian and optimistic dual methods for equality-constrained optimization by proposing an additive hybrid framework that unifies matrix-valued augmentation and optimistic correction as distinct decompositions of a common correction matrix. By adaptively selecting the optimal splitting and stepsize through local spectral weighting, the method yields, for the first time, a closed-form hybrid update rule that jointly balances primal curvature and dual memory scale within a finite number of steps. Theoretical analysis reveals the equivalence and design flexibility between the two mechanisms, while experiments demonstrate that the proposed approach significantly outperforms individual strategies on nonlinear equality-constrained problems, achieving performance close to grid-search optimality and matching state-of-the-art first-order primal-dual algorithms under moderate ill-conditioning.

augmented Lagrangianconstrained optimizationfeasibility

Cautious Weight Decay

Oct 14, 2025
LC
Lizhang Chen
🏛️ University of Texas at Austin | Google

Standard weight decay uniformly penalizes all parameters, interfering with the optimizer’s original objective and hindering convergence to local optima of the unmodified loss. To address this, we propose Sign-Aligned Weight Decay (SAWD): it applies decay only to parameter coordinates whose signs align with those of the corresponding gradients, thereby preserving the original loss function exactly. SAWD further introduces a bilevel optimization mechanism that identifies locally Pareto-stationary points of the unmodified objective—without introducing additional hyperparameters. Inspired by sliding-mode control, SAWD employs sign-comparison logic for selective regularization and is compatible with mainstream optimizers including AdamW and Lion. Empirically, on pretraining billion-parameter language models and ImageNet classification, SAWD consistently reduces final loss and improves accuracy.

Improves model performance without requiring new hyperparameters or tuningModifies weight decay to apply only when signs match optimizer updatesPreserves original loss function while enabling sliding-mode optimization behavior

Latest Papers

What's happening recently
View more

This work addresses the high computational cost of exactly computing all optimal solutions that satisfy the priority structure specified in a rulebook for multi-objective robotic planning. To overcome this challenge, the paper introduces, for the first time, the concept of ε-rule-dominance as an approximation criterion and proposes the RA*pex algorithm, which efficiently generates a compact set of near-optimal solutions via best-first search. By integrating dimensionality reduction, hierarchical closed-set maintenance, and residual rule-set dominance checks, RA*pex significantly reduces computational complexity while strictly preserving the semantic hierarchy of the rulebook. Experimental results demonstrate that RA*pex outperforms existing algorithms by over two orders of magnitude in speed, with theoretical guarantees that every solution in the returned set ε-rule-dominates some true rulebook-optimal solution.

approximate optimizationcomputational complexitymulti-objective search

This work addresses the lack of machine-verifiable formalizations of line search methods in nonlinear optimization, which has hindered algorithmic reliability. Within the Lean 4 theorem prover, it presents the first systematic formalization of several classical line search criteria—including Armijo, Goldstein, Wolfe, and their nonmonotone variants—alongside rigorous definitions of gradient descent, descent directions, and backtracking step-size selection. The study fully verifies the Zoutendijk convergence theorem within this framework, thereby establishing a comprehensive formal foundation for line search theory. This contribution significantly enhances the verifiability and trustworthiness of nonlinear optimization algorithms through mechanized mathematical reasoning.

convergenceformalizationline search

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This work addresses the vulnerability of large language models to prompt injection attacks and their difficulty in resolving conflicting instructions, stemming from the absence of a structured priority mechanism for multi-source directives. The study formalizes, for the first time, the k-level instruction hierarchy problem and introduces a five-tier privilege framework. To enforce hierarchical compliance, the authors propose Gravitational Weighting Direct Preference Optimization (GW-DPO), which integrates hierarchical delimiters and instruction segment embeddings with a bilateral scheduling strategy that dynamically weights the severity of violations. Experiments on Llama-3.1-8B-Instruct demonstrate that GW-DPO achieves a significant improvement in macro-level adherence to instruction hierarchies while maintaining an over-rejection rate only half that of standard DPO, thereby yielding a Pareto improvement over existing approaches.

conflicting instructionsinstruction hierarchyLLM alignment

Hot Scholars

HS

Haijian Sun

Assistant Professor of ECE, University of Georgia
5G and BeyondV2XmmWave SensingMachine Learning on Edge
YS

Yi-Shuai Niu

Beijing Institute of Mathematical Sciences and Applications (BIMSA)
OptimizationMachine LearningHigh-Performance Computing
RQ

Rose Qingyang Hu

IEEE Fellow, AAAS Fellow, Virginia Tech
Wireless communications and networksMobile edge computingInternet of ThingsAI/ML
ZA

Zhenlin An

Assistant Professor, University of Georgia
Wireless SystemInternet of ThingsSensingMachine Learning
YP

Yiyang Peng

Imperial College London
Wireless Communications