Score
Designs and implements neural-network architectures that convert iterative optimization algorithms into fixed computational graphs by unrolling a finite number of algorithm iterations into learned layers; builds models that parameterize update steps (e.g., alternating variable updates or augmented Lagrangian iterations), incorporate learned proximal operators or denoisers as priors, and enable end-to-end training from data. Analyzes the resulting models' stability, convergence behavior, and interpretability, and tunes layer-wise parameterizations and training regimes to balance algorithmic structure and data-driven flexibility.
To address the low training efficiency, poor generalization, and strong hyperparameter sensitivity of neural networks across varying scales, this paper proposes a scale-invariant adaptive optimization framework. The method unifies adaptive optimization, second-order information approximation, learning-rate scaling invariance, and gradient compression, thereby decoupling optimization from model size and hardware configuration. Its core innovation lies in a scale-robust update paradigm that ensures stable optimization dynamics under variations in parameter count, batch size, and device count. Extensive experiments across diverse architectures—including MLPs, CNNs, and Transformers—and benchmarks—including CIFAR-10/100, ImageNet, and WikiText—demonstrate that the framework achieves 1.3–2.1× speedup over baseline optimizers, improved convergence stability, significantly reduced hyperparameter sensitivity, and eliminates the need for scale-specific hyperparameter tuning.
In recent years, there has been growing interest in understanding neural architectures'ability to learn to execute discrete algorithms, a line of work often referred to as neural algorithmic reasoning. The goal is to integrate algorithmic reasoning capabilities into larger neural pipelines. Many such architectures are based on (message-passing) graph neural networks (MPNNs), owing to their permutation equivariance and ability to deal with sparsity and variable-sized inputs. However, existing work is either largely empirical and lacks formal guarantees or it focuses solely on expressivity, leaving open the question of when and how such architectures generalize beyond a finite training set. In this work, we propose a general theoretical framework that characterizes the sufficient conditions under which MPNNs can learn an algorithm from a training set of small instances and provably approximate its behavior on inputs of arbitrary size. Our framework applies to a broad class of algorithms, including single-source shortest paths, minimum spanning trees, and general dynamic programming problems, such as the $0$-$1$ knapsack problem. In addition, we establish impossibility results for a wide range of algorithmic tasks, showing that standard MPNNs cannot learn them, and we derive more expressive MPNN-like architectures that overcome these limitations. Finally, we refine our analysis for the Bellman-Ford algorithm, yielding a substantially smaller required training set and significantly extending the recent work of Nerem et al. [2025] by allowing for a differentiable regularization loss. Empirical results largely support our theoretical findings.
Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.
Existing preconditioners for large-scale sparse linear systems suffer from low efficiency and poor generalization. Method: This paper proposes a novel learnable preconditioner that integrates algebraic preconditioning with graph neural networks (GNNs). It initializes the GNN with a classical ILU-type preconditioner and introduces a differentiable, condition-number-based loss function to explicitly optimize spectral properties during training. Additionally, it incorporates sparse structural priors and parameterized PDE modeling to ensure physical consistency and computational tractability. Results: On benchmark discretized parametric PDE systems, the method reduces iterative solver iterations by 30–50% compared to ILU and state-of-the-art neural preconditioners, achieves significantly improved condition numbers, and incurs only modest inference overhead. This work overcomes key limitations of purely data-driven and purely sparse-GNN-based preconditioners, establishing a new paradigm for interpretable, efficient, and generalizable learning in numerical linear algebra.
Traditional algebraic preconditioners (e.g., ILU, AMG) suffer from failure on ill-conditioned large-scale sparse linear systems, exhibit unpredictable and expensive setup costs, and often rely on problem-specific physical priors. Method: We propose the first end-to-end differentiable, general-purpose preconditioner based on graph neural networks (GNNs). It encodes sparse matrices as graphs without requiring underlying physical knowledge, enabling strong generalization across diverse problem domains. Contribution/Results: The GNN-based preconditioner achieves highly predictable and significantly accelerated setup times compared to ILU and AMG. Integrated tightly with Krylov subspace methods (e.g., GMRES), it reduces iteration counts relative to inner-outer GMRES. Evaluated on 800+ real-world matrices spanning PDEs, economics, statistics, and graph learning, our approach consistently improves both solver efficiency and robustness—demonstrating superior scalability, generality, and practical applicability for large-scale sparse linear systems.
This work addresses the challenge of efficiently solving NP-hard Ising and Max-Cut problems, where conventional methods struggle due to the complex, non-convex energy landscapes. The authors propose a data-driven iterative dynamical system that parameterizes spin update rules via a shared node-level multilayer perceptron and trains it using zeroth-order optimization to circumvent gradient instability associated with backpropagation. Remarkably, with an extremely low number of parameters, the learned dynamics automatically exhibit momentum-like behavior and time-varying scheduling mechanisms, substantially enhancing search efficiency. Evaluated on standard Ising and combinatorial optimization benchmarks, the method achieves solution quality and convergence speed comparable to state-of-the-art learning-based approaches and classical Ising machine heuristics.
This work addresses the challenge of characterizing the highly complex loss landscape in large language model (LLM) pretraining, where existing theories struggle to balance analytical tractability with accurate dynamic prediction. By performing Taylor expansions of both the model and loss function at mid-training, the authors construct a local quadratic approximation and combine it with Lanczos quadrature and Hessian spectral estimation. For the first time, they validate this approach on a 150M-parameter LLM trained on 3B tokens, demonstrating predictive accuracy over a training window spanning 10% of total steps. Their analysis reveals that the quadratic model faithfully captures optimization trajectories, that the tail structure of the Hessian spectrum is strongly influenced by batch size, preconditioning, and training stage, and that optimization typically resides in a stochastic edge-of-stability regime dictated by batch size—uncovering a deep connection between local stability and hyperparameter choice.
This study addresses the limitation of plug-and-play methods, where directly substituting proximal operators compromises variational interpretability and precludes convergence guarantees. To overcome this, we propose the Learned Proximal Network (LPN), which leverages architectural design to ensure that the denoiser strictly corresponds to the proximal operator of a regularizer. Furthermore, we extend the theoretical framework to a broader class of activation functions, analytically characterize the mean-induced regularization mechanism, and develop an operator scaling technique with provable convergence guarantees. Consequently, this work restores both the variational interpretation and convergence properties of the algorithm while maintaining state-of-the-art reconstruction quality, thereby providing rigorous theoretical foundations for plug-and-play approaches.
Stochastic Gradient Descent (SGD) and its variants lack rigorous theoretical foundations in over-parameterized neural networks, suffering from inefficient training and poor interpretability. Method: This paper proposes a principle-driven guided descent framework that unifies, for the first time, curvature-aware second-order approximations, layer-adaptive preconditioning (calibrated via condition number), and a dynamically parameterized maximum-update learning rate mechanism. It systematically elucidates the synergistic interplay between this framework and exponential moving average (EMA) as well as learning rate scheduling. Contribution/Results: The method achieves both scalability and theoretical interpretability while preserving training stability and significantly accelerating convergence—reducing large-model training time by an order of magnitude. Moreover, it enhances discriminative feature learning, simultaneously improving generalization performance and output consistency.
This work addresses the lack of a unified theoretical foundation for learned iterative networks in computational imaging and inverse problems. We propose a continuous-domain reconstruction framework grounded in operator learning, which explicitly decouples *how to compute* (algorithmic architecture) from *what to compute* (the reconstruction operator), thereby bridging the theoretical gap between classical optimization-based methods and data-driven models. Methodologically, we integrate variational unfolding, operator modeling in function spaces, deep neural network parameterization, and end-to-end training into a single coherent framework—yielding a learnable, interpretable, and generalizable reconstruction operator. Our approach unifies major classes of learned iterative methods under a common theoretical umbrella. Extensive numerical experiments validate its effectiveness. The framework establishes a new paradigm for designing reconstruction networks that simultaneously offer rigorous theoretical guarantees and strong practical performance.