weight-space interpolation

Designs and evaluates procedures that combine trained model parameter vectors by interpolating their weights (e.g., linear or affine mixing between a base and a fine-tuned model) to produce a single parameter set with controlled tradeoffs between specialization and generalization. Analyzes interpolation strategies, weight-space geometry, and their effects on metrics such as accuracy, robustness, calibration, and performance across evaluation benchmarks.

weight-spaceinterpolation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work addresses the long-standing neglect of the structural and knowledge-rich properties inherent in neural network weights, as well as the lack of systematic investigation into weight space within conventional deep learning. To bridge this gap, the paper introduces Weight Space Learning (WSL), a unified framework that establishes the first comprehensive taxonomy for studying weight space through three core dimensions: understanding (geometric structure and symmetries), representation (model embeddings), and generation (hypernetworks and generative models). By integrating previously fragmented research efforts, WSL reveals the potential of weights as a learnable, structured domain and enables advances in diverse applications—including model retrieval, continual learning, federated learning, neural architecture search, and data-free reconstruction. The authors further support community progress by releasing an open-source repository dedicated to weight space research.

Knowledge TransferModel AnalysisNeural Network Weights

Must-Read Papers

Most classic and influential ideas
View more

LAMP: Data-Efficient Linear Affine Weight-Space Models for Parameter-Controlled 3D Shape Generation and Extrapolation

Oct 25, 2025
GN
Ghadi Nehme
🏛️ Massachusetts Institute of Technology | Toyota Research Institute

To address the strong data dependency, poor controllability, and limited generalization in parametric 3D shape generation, this paper proposes a data-efficient, controllable, and interpretable generative framework. Methodologically, it performs linear affine mixing in an aligned neural weight space, integrates SDF decoder overfitting with parameter-constrained optimization, and introduces a linear mismatch safety metric to ensure geometric validity. The approach achieves high-fidelity interpolation and safe extrapolation using only ~100 training samples—enabling, for the first time, full-range (100%) parameter-space extrapolation. Experiments on DrivAerNet++ and BlendedNet demonstrate substantial improvements over conditional autoencoders and Deep Neural Interpolation (DNI), particularly in data efficiency, precise parametric control, and physics-informed performance optimization.

Enabling controlled interpolation and safe extrapolation beyond training rangesGenerating high-fidelity 3D shapes with parameter constraintsOvercoming data inefficiency and limited generalization in 3D generation

This work investigates model collapse in overparameterized linear regression arising from iterative mixing of ground-truth labels with model-generated synthetic labels. We propose an iterative learning framework incorporating a tunable label-mixing ratio and conduct rigorous asymptotic analysis of the generalization error for both minimum-ℓ₂-norm interpolation and ridge regression. Our key theoretical finding is that, under minimum-ℓ₂-norm interpolation, the optimal proportion of ground-truth labels converges to the inverse golden ratio (≈0.618), a phenomenon governed by the geometric structure of the data covariance spectrum; this result formally establishes that model collapse is avoidable. We further derive exact asymptotic generalization error expressions, proving that the optimal mixing ratio is always at least 1/2—i.e., ground-truth data must dominate. Extensive simulations corroborate the theory, providing an interpretable, principled foundation for robust self-training.

Analyzing spectral geometry impact on preventing model degradationDeriving optimal mixing ratios for interpolation and ridge regressionStudying model collapse in overparameterized linear regression

This study addresses the prohibitive computational cost of optimizing large language model merging coefficients, which typically relies on expensive benchmark evaluations. We propose a quadratic surrogate modeling approach grounded in mixture design methodology. Specifically, classical mixture designs are employed over the coefficient simplex to determine a minimal set of measurement points, replacing random sampling with structured evaluation. A quadratic surrogate model is then constructed and interpolated over the simplex to predict multi-expert mixture performance while minimizing evaluation overhead. This method accurately predicts outcomes for unseen coefficient combinations, achieving competitive model merging performance under limited computational budgets and significantly reducing the overall cost of merging optimization.

Evaluation BudgetLarge Language ModelsMerging Coefficients Optimization

Tailoring Mixup to Data for Calibration

Nov 02, 2023
QB
Quentin Bouniot
🏛️ Télécom Paris | Institut Polytechnique de Paris

Mixup’s blind interpolation often yields synthetic samples deviating from the underlying data manifold, degrading model calibration. We observe that interpolation distance correlates positively with mislabeling risk. To address this, we propose a similarity-driven adaptive Mixup framework that dynamically modulates the Beta distribution parameter based on feature-space distances, prioritizing interpolation between nearby samples. Our method preserves augmentation diversity while significantly mitigating manifold mismatch. Extensive experiments across multi-class classification and regression tasks demonstrate an average 35% reduction in Expected Calibration Error (ECE), consistent improvements in Brier Score, enhanced accuracy, and 2.1× faster training convergence. This work establishes the first coupling of sample similarity modeling with interpolation coefficient distribution control, introducing a novel paradigm for calibration-aware data augmentation.

Addresses calibration issues in Mixup data augmentationImproves model predictive performance and calibration efficiencyReduces wrong label likelihood by adjusting interpolation coefficients

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified theoretical foundation in existing model merging approaches and the opacity of hyperparameters in open-source fine-tuned models, which together hinder the predictability of merged model performance. Leveraging L2-stability theory, the study establishes the first unified generalization framework to systematically analyze the generalization capability of merged heterogeneous expert models and proposes actionable fine-tuning strategies to enhance mergeability. Through parameter-space merging, derivation of generalization bounds, and large-scale vision experiments on ResNet and ViT architectures, the authors validate the critical influence of hyperparameters on merging performance across 20 and 8 tasks, respectively. Theoretical predictions align closely with empirical results, significantly improving the predictability and effectiveness of model merging.

fine-tuned modelsgeneralizationheterogeneous hyperparameters

Existing parameter reparameterization methods are often confined to a single objective—either parameter-efficient fine-tuning or model compression—making it challenging to simultaneously address both demands under resource constraints. This work proposes CRISP, a unified framework that jointly achieves model compression and parameter-efficient fine-tuning within a single architecture. CRISP decomposes pre-trained weights into shared base matrices and lightweight mixture coefficients, enhanced by cross-layer base sharing and an interpolation-based gated coefficient recombination mechanism. Requiring fewer than 200 trainable parameters, CRISP outperforms existing approaches by 1% in joint compression and fine-tuning tasks, surpasses state-of-the-art methods by up to 1.5% in pure parameter-efficient fine-tuning, and achieves a consistent 4–5% improvement in overall dual-task performance.

Edge DeploymentModel CompressionNeural Network Compression

Deep Parameter Interpolation for Scalar Conditioning

Nov 25, 2025
CY
Chicago Y. Park
🏛️ WashU | Los Alamos National Laboratory | UW–Madison

Existing generative models struggle to efficiently fuse heterogeneous conditions—such as high-dimensional image features and scalar-valued time/noise levels—due to architectural coupling or redundant encoding in mainstream approaches. To address this, we propose Deep Parameter Interpolation (DPI), a lightweight, plug-and-play conditioning mechanism that dynamically interpolates between two pretrained parameter sets within the neural network’s layer-wise parameter space, guided by a learnable scalar encoding. DPI is architecture-agnostic: it requires no modification to the backbone network and seamlessly integrates with diverse generative paradigms, including diffusion models and flow matching. Experiments demonstrate that DPI improves denoising accuracy and generation quality while maintaining computational overhead comparable to baseline methods. Its generalizability is validated across multiple generative tasks, confirming robust performance without sacrificing efficiency or flexibility.

Enabling neural networks to process scalar inputs alongside high-dimensional dataImproving generative model performance through dynamic parameter interpolationOvercoming architecture limitations in scalar-vector information integration

This work investigates model size interpolation under zero-shot settings via layer patching, aiming to construct intermediate-scale models whose performance lies between that of a teacher and a student model. The authors formulate layer patching for the first time as a shortest-path optimization problem on a directed acyclic graph, revealing the critical impact of patching order on interpolation efficacy. They propose KLPatch, a greedy algorithm guided by KL divergence, to determine an effective patching sequence. Experimental results demonstrate that even a simple head-to-tail patching strategy yields competitive performance, while KLPatch further enhances interpolation quality. This study provides both theoretical insights and a practical method for efficiently constructing near-optimal intermediate models without additional training.

language modelslayer patchingmodel size interpolation

This work addresses the lack of provable generalization guarantees for multidimensional hyperparameter tuning in complex, non-smooth spaces. It proposes the first data-driven framework for such settings, establishing generalization bounds under mild assumptions by integrating structured loss with validation loss. Leveraging tools from real algebraic geometry, the analysis characterizes the complexity of semi-algebraic function classes, yielding tighter and more broadly applicable generalization bounds. The framework’s effectiveness and learnability are demonstrated on models such as weighted group Lasso and weighted fused Lasso, offering both theoretical foundations and practical methodologies for multidimensional hyperparameter optimization.

data-drivengeneralization guaranteeshyperparameter tuning

Hot Scholars

JL

Junlin Li

ByteDance Inc. - Georgia Institute of Technology - Tsinghua University
Video Compression and ProcessingVideo StreamingMachine LearningAI
HK

Ho-Kin Tang

Harbin Institute of Technology (Shenzhen)
strongly correlated systemsquantum Monte Carlooptimization algorithms
DH

Daojing He

School of Computer Science and Engineering, South China University of Technology
Network and Information security
GD

Guodong Du

Yanshan University, China
Machine learningData MiningAI in MedicineHealth Informatics
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing