Score
Designs, implements, and analyzes training procedures and parameterizations for hypernetworks—models that generate the weights of another model—including learning hypernetwork weights, lifted parameterizations that stochastically emit layer weights, and mechanisms to enforce or relax constraints on emitted weights (for example, producing non‑negative inter-layer weights). This skill covers constructing emission and sampling schemes, choosing losses and optimization strategies to soften the optimization landscape, and implementing constraint-handling techniques that avoid hard projections or softplus-related optimization stalls to improve convergence.
This work addresses the long-standing neglect of the structural and knowledge-rich properties inherent in neural network weights, as well as the lack of systematic investigation into weight space within conventional deep learning. To bridge this gap, the paper introduces Weight Space Learning (WSL), a unified framework that establishes the first comprehensive taxonomy for studying weight space through three core dimensions: understanding (geometric structure and symmetries), representation (model embeddings), and generation (hypernetworks and generative models). By integrating previously fragmented research efforts, WSL reveals the potential of weights as a learnable, structured domain and enables advances in diverse applications—including model retrieval, continual learning, federated learning, neural architecture search, and data-free reconstruction. The authors further support community progress by releasing an open-source repository dedicated to weight space research.
To address poor generalization in parametric dynamical systems caused by parameter variability, this paper proposes PHLieNet—a physics-informed hypernetwork framework. PHLieNet learns a nonlinear embedding of the parameter space via a hypernetwork and dynamically generates weights for a Lie group–based physical propagation network, enabling adaptive prediction across diverse parameter configurations. Crucially, it performs interpolation in model weight space—not observation space—thereby supporting smooth cross-parameter transfer and robust extrapolation/interpolation. The framework unifies parameter-conditioned weight generation, nonlinear parameter embedding learning, and sequential modeling to construct a tunable foundational dynamics network. Evaluated on canonical parametric systems—including Lorenz-96 and Kuramoto–Sivashinsky equations—PHLieNet achieves state-of-the-art or competitive performance in both short-term forecasting accuracy and long-term statistical fidelity (e.g., attractor structure preservation).
This work addresses the bottleneck in hypernetwork training that relies on per-sample ground-truth weights. We propose HyperNet Field, a novel paradigm that models task-network weights as an input-conditioned continuous neural field, implicitly learning their optimization trajectory—not just the final converged state. Our method enables end-to-end training solely via gradient consistency constraints, eliminating the need for any sample-level weight supervision. Key technical components include neural field parameterization, implicit trajectory modeling, and gradient matching. The framework unifies support for personalized image generation and single-image or single-point-cloud-driven 3D reconstruction. Experiments demonstrate competitive performance across diverse tasks, establishing the first hypernetwork training approach that operates entirely without sample-level weight supervision.
This study addresses the limitation of few-shot learning approaches that typically rely on task-specific training, making it difficult to directly acquire specialized model parameters from limited demonstrations. To overcome this, we propose a hypernetwork-based context weight generation mechanism that synthesizes micro-expert model weights on the fly without requiring explicit task identifiers, thereby enabling rapid compilation and execution for few-shot tasks. Experimental evaluations on the ARC-1D benchmark validate the feasibility of dynamic weight generation and demonstrate that structured weight spaces effectively support compositional generalization. Consequently, the proposed approach achieves functional generalization capabilities that extend beyond the training distribution, offering a promising paradigm for adaptive few-shot learning.
本文针对流数据中超参数优化问题,提出四种边界约束处理策略,通过实验验证其优于现有方法。
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
本文提出了一种通过自动微分优化拓扑结构和超参数的方法,解决了拓扑优化中超参数调优的问题,且该方法可扩展至数千个超参数。
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work addresses the limited generality of conventional hypernetworks, which are typically constrained to specific architectures and tasks. The authors propose a Universal HyperNetwork (UHN) that employs a fixed-architecture generator to uniformly predict the weights of arbitrary target models based on deterministic encodings of parameters, architectural specifications, and task descriptors. UHN is the first framework to enable a single, fixed hypernetwork to generate models across heterogeneous architectures and diverse tasks, while demonstrating stable three-level recursive generation. Experimental results show that UHN achieves performance comparable to directly trained models across a range of domains—including vision, graph neural networks, text processing, and symbolic regression—significantly enhancing generalization across multiple models and multitask learning capabilities.
研究通过建立基准优化器与超球优化器的等价性,提出HyperTransfer方法,仅使用初始化和学习率调度即可复制基准优化器的动力学特性。
This work addresses the inefficiency of global optimization when standard neural network surrogates are embedded into mixed-integer linear programs (MILPs), a challenge stemming from the lack of control over their structural properties. The authors propose a novel differentiable regularizer that, for the first time, approximates the full gradient of the LP relaxation gap with respect to network parameters, enabling direct optimization of key structural attributes such as big-M constants, the number of unstable neurons, and the LP relaxation gap itself. Built upon ReLU networks and MILP formulations, the method leverages gradients from LP dual variables and requires no custom automatic differentiation. Experiments demonstrate up to four orders of magnitude reduction in MILP solve time on nonconvex benchmark functions and two-stage stochastic programming problems, all while preserving predictive accuracy.