bilevel neural architecture search

Designs and implements neural architecture search systems that formulate architecture selection as a bilevel optimization problem in which the outer problem optimizes architecture variables (e.g., topology or operator choices) to minimize validation loss while the inner problem optimizes network weights to minimize training loss. Builds and analyzes algorithms—exact, approximate, or differentiable—that coordinate updates between architecture and weights and evaluate their interaction, convergence, and computational trade-offs.

bilevelneuralarchitecturesearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a bilevel neural architecture search (NAS) framework grounded in auxiliary mathematical programming, formulating NAS as a bilevel optimization problem that jointly optimizes outer-level architecture parameters and inner-level network weights. By explicitly incorporating second-order derivative information of the training loss, the method enables synchronous updates of architecture and model parameters while ensuring local optimality of the inner-level solution. The approach systematically integrates bilevel optimization theory with mathematical programming techniques to enhance both search efficiency and final model performance. Experimental results demonstrate that the proposed framework significantly outperforms conventional sampling-based NAS methods in terms of both accuracy and computational efficiency.

Architecture ParametersBilevel OptimizationHierarchical Optimization

CR-BLEA: Contrastive Ranking for Adaptive Resource Allocation in Bilevel Evolutionary Algorithms

Jun 03, 2025
DX
Dejun Xu
🏛️ Xiamen University | Sichuan University

In bilevel optimization, evaluating upper-level solutions requires repeated resolution of the lower-level problem, leading to substantial computational redundancy. This paper proposes an adaptive resource allocation framework that models the upper–lower-level solution relationship online and dynamically identifies and prioritizes high-potential lower-level tasks. Its core contribution is a reference-based ranking strategy driven by a contrastive ranking network, enabling population-quality-aware adaptive resampling—overcoming limitations of conventional static or heuristic allocation schemes. Evaluated across five state-of-the-art bilevel evolutionary algorithms, the framework reduces function evaluations by 37.2% on average while preserving or improving solution accuracy. It exhibits strong generalizability and plug-and-play compatibility.

Adaptively allocates resources to promising lower-level tasksImproves computational efficiency without sacrificing solution accuracyReduces redundant lower-level task evaluations in bilevel EAs

On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis

Jan 02, 2023
LC
Le‐Yu Chen
🏛️ Tsinghua University | Shanghai Qizhi Institute | Shanghai AI Lab

This work investigates the theoretical hardness and algorithmic efficiency of finding stationary points of the hyperobjective in nonconvex–convex and nonconvex–nonconvex bilevel optimization, under the Polyak–Łojasiewicz (PL) condition—rather than strong convexity—on the lower-level objective. We first establish an impossibility result: for zero-respecting algorithms, computing a hyperstationary point is fundamentally intractable in the nonconvex–convex setting. Under the PL condition, we break the reliance on strong convexity and derive tighter hypergradient convergence complexity bounds. We propose a novel analytical framework unifying implicit function differentiation, hypergradient estimation, and first-order optimization. This yields complexity guarantees of $ ilde{mathcal{O}}(varepsilon^{-2})$, $ ilde{mathcal{O}}(varepsilon^{-4})$, and $ ilde{mathcal{O}}(varepsilon^{-6})$ for deterministic, partially stochastic, and fully stochastic settings, respectively—substantially improving upon existing nonconvex bilevel optimization methods.

Analyzing hyper-objective optimization complexity without strong convexityDeveloping efficient algorithms for PL-condition nonconvex-nonconvex problemsProviding hardness results for nonconvex-convex bilevel optimization

This paper addresses the computational intractability arising from millions of independent followers (e.g., cyclists) in large-scale bilevel and two-stage stochastic programming, using bicycle infrastructure planning as a motivating application. We propose an integrated optimization framework combining Monte Carlo sampling, graph neural network (GNN)-based representation learning, and ensemble regression modeling. Our contributions are threefold: (1) the first follower sampling and embedded machine learning estimation method with provable theoretical optimality guarantees; (2) a bilevel-aware follower representation learning mechanism leveraging GNNs to encode spatial and behavioral heterogeneity; and (3) scalability to real-world city-scale networks with up to one million nodes. Evaluated on Toronto’s urban transportation system, our approach improves cycling accessibility by 19.2%—equivalent to $18 million in societal cost savings—while significantly outperforming baseline methods in decision quality. The framework has been deployed in operational infrastructure planning.

Enhance decision quality using machine learningImprove cycling network design efficiencyOptimize large bilevel and stochastic programs

Latest Papers

What's happening recently
View more

This work addresses the challenge of catastrophic forgetting in continual learning, where static neural architectures struggle to adapt to shifting data distributions. The authors propose a unified framework that jointly models network architecture and weights within a Sobolev space, employing bilevel optimization to simultaneously learn both components. To handle parameter dimension mismatches during architectural evolution, they introduce a low-rank knowledge transfer mechanism. Theoretically, they provide the first rigorous proof that weight-only optimization is insufficient to mitigate forgetting, thereby establishing a formal foundation for the co-adaptation of architecture and weights, along with a derivative-free direct search algorithm. Experiments across diverse networks and tasks demonstrate performance improvements of up to two orders of magnitude, significantly alleviating forgetting and enhancing robustness to noise.

catastrophic forgettingcontinual learningdistribution shift

This work addresses the high computational cost of black-box bilevel optimization, particularly in multimodal settings with strong variable interactions where solving nested problems is notoriously difficult. The authors propose an efficient framework that, for the first time, incorporates a ranking-based approximation of the upper-level value function into bilevel optimization. By leveraging the invariance of rank-based evolutionary algorithms to monotonic transformations, the method directly models the ordinal relationships of the upper-level objective, thereby circumventing the need to fully converge the lower-level optimization at every iteration. Using CMA-ES as the continuous optimizer, the approach constructs a rank-invariant surrogate model that substantially reduces computational overhead. Empirical results on standard benchmarks demonstrate superior performance and the ability to solve complex bilevel problems previously intractable to existing methods.

bilevel optimizationblack-box optimizationcomputational cost

Hot Scholars

LW

Longfeng Wu

Virginia Tech
Machine LearningRecommender SystemsGraph Mining
TS

Tim Schlippe

Silicon Surfer & IU International University of Applied Sciences
Artificial IntelligenceSpeech ProcessingNatural Language Processing
LZ

Lecheng Zheng

University of Illinois at Urbana-Champaign
Heterogeneous LearningGraph MiningMulti-modal LearningAnomaly Detection
SH

Sitao Huang

Assistant Professor of EECS, University of California Irvine
Hardware AccelerationHigh-Level SynthesisFPGAParallel Computing