train hypernetworks

Designs, implements, and analyzes training procedures and parameterizations for hypernetworks—models that generate the weights of another model—including learning hypernetwork weights, lifted parameterizations that stochastically emit layer weights, and mechanisms to enforce or relax constraints on emitted weights (for example, producing non‑negative inter-layer weights). This skill covers constructing emission and sampling schemes, choosing losses and optimization strategies to soften the optimization landscape, and implementing constraint-handling techniques that avoid hard projections or softplus-related optimization stalls to improve convergence.

trainhypernetworks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address poor generalization in parametric dynamical systems caused by parameter variability, this paper proposes PHLieNet—a physics-informed hypernetwork framework. PHLieNet learns a nonlinear embedding of the parameter space via a hypernetwork and dynamically generates weights for a Lie group–based physical propagation network, enabling adaptive prediction across diverse parameter configurations. Crucially, it performs interpolation in model weight space—not observation space—thereby supporting smooth cross-parameter transfer and robust extrapolation/interpolation. The framework unifies parameter-conditioned weight generation, nonlinear parameter embedding learning, and sequential modeling to construct a tunable foundational dynamics network. Evaluated on canonical parametric systems—including Lorenz-96 and Kuramoto–Sivashinsky equations—PHLieNet achieves state-of-the-art or competitive performance in both short-term forecasting accuracy and long-term statistical fidelity (e.g., attractor structure preservation).

Enabling adaptive forecasting for unseen system dynamicsGeneralizing models across different parameter regimesHandling parametric variability in dynamical systems modeling

HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories

Dec 22, 2024
EH
Eric Hedlin
🏛️ University of British Columbia | Qualcomm AI Research

This work addresses the bottleneck in hypernetwork training that relies on per-sample ground-truth weights. We propose HyperNet Field, a novel paradigm that models task-network weights as an input-conditioned continuous neural field, implicitly learning their optimization trajectory—not just the final converged state. Our method enables end-to-end training solely via gradient consistency constraints, eliminating the need for any sample-level weight supervision. Key technical components include neural field parameterization, implicit trajectory modeling, and gradient matching. The framework unifies support for personalized image generation and single-image or single-point-cloud-driven 3D reconstruction. Experiments demonstrate competitive performance across diverse tasks, establishing the first hypernetwork training approach that operates entirely without sample-level weight supervision.

Enabling efficient large model adaptation and generative modelingLearning weight trajectories instead of converged statesTraining hypernetworks without per-sample ground truth weights

This study addresses the limitation of few-shot learning approaches that typically rely on task-specific training, making it difficult to directly acquire specialized model parameters from limited demonstrations. To overcome this, we propose a hypernetwork-based context weight generation mechanism that synthesizes micro-expert model weights on the fly without requiring explicit task identifiers, thereby enabling rapid compilation and execution for few-shot tasks. Experimental evaluations on the ARC-1D benchmark validate the feasibility of dynamic weight generation and demonstrate that structured weight spaces effectively support compositional generalization. Consequently, the proposed approach achieves functional generalization capabilities that extend beyond the training distribution, offering a promising paradigm for adaptive few-shot learning.

ARC-1DFew-shot LearningHypernetwork

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This work addresses the limited generality of conventional hypernetworks, which are typically constrained to specific architectures and tasks. The authors propose a Universal HyperNetwork (UHN) that employs a fixed-architecture generator to uniformly predict the weights of arbitrary target models based on deterministic encodings of parameters, architectural specifications, and task descriptors. UHN is the first framework to enable a single, fixed hypernetwork to generate models across heterogeneous architectures and diverse tasks, while demonstrating stable three-level recursive generation. Experimental results show that UHN achieves performance comparable to directly trained models across a range of domains—including vision, graph neural networks, text processing, and symbolic regression—significantly enhancing generalization across multiple models and multitask learning capabilities.

hypernetworksmodel architecturemulti-task learning

This work addresses the inefficiency of global optimization when standard neural network surrogates are embedded into mixed-integer linear programs (MILPs), a challenge stemming from the lack of control over their structural properties. The authors propose a novel differentiable regularizer that, for the first time, approximates the full gradient of the LP relaxation gap with respect to network parameters, enabling direct optimization of key structural attributes such as big-M constants, the number of unstable neurons, and the LP relaxation gap itself. Built upon ReLU networks and MILP formulations, the method leverages gradients from LP dual variables and requires no custom automatic differentiation. Experiments demonstrate up to four orders of magnitude reduction in MILP solve time on nonconvex benchmark functions and two-stage stochastic programming problems, all while preserving predictive accuracy.

big-M constantsLP relaxationMILP tractability

Hot Scholars

BK

Bostan Khan

Mälardalen University
Computer VisionDeep LearningMachine Learning
MD

Masoud Daneshtalab

Professor and Head of DeepHERO Lab.
Deep LearningHeterogeneous and Dependable ComputingInterconnection Networks
YY

Yong Yu

Materials Engineer
Polymer matrix compositeadhesivemodelingtest development
SG

Shangqian Gao

Florida State University
Computer VisionNatural Lanugage ProcessingMachine Learning
ZX

Zhitong Xiong

Technical Univercity of Munich
Deep LearningRemote SensingComputer Vision