Score
Designs and evaluates procedures that combine trained model parameter vectors by interpolating their weights (e.g., linear or affine mixing between a base and a fine-tuned model) to produce a single parameter set with controlled tradeoffs between specialization and generalization. Analyzes interpolation strategies, weight-space geometry, and their effects on metrics such as accuracy, robustness, calibration, and performance across evaluation benchmarks.
This work addresses fundamental challenges in model merging—including the absence of a unified taxonomy, terminological inconsistency, incomparable methodologies, and difficulties in multi-task fusion under data-unavailable scenarios. We propose the first three-tiered classification paradigm encompassing weight-space fusion, gradient alignment, and task disentanglement. We establish a cross-method reproducible evaluation benchmark and formally define and distinguish the applicability boundaries of “data-agnostic” versus “data-aware” merging. By unifying the theoretical formulations of over 20 state-of-the-art methods—via spectral analysis, normalization sensitivity diagnosis, and task vector geometric modeling—we identify three root causes of merging failure: directional conflict, scale mismatch, and task entanglement. Our framework provides systematic theoretical foundations and principled design guidelines for efficient, lightweight, and interpretable model fusion.
This work addresses the long-standing neglect of the structural and knowledge-rich properties inherent in neural network weights, as well as the lack of systematic investigation into weight space within conventional deep learning. To bridge this gap, the paper introduces Weight Space Learning (WSL), a unified framework that establishes the first comprehensive taxonomy for studying weight space through three core dimensions: understanding (geometric structure and symmetries), representation (model embeddings), and generation (hypernetworks and generative models). By integrating previously fragmented research efforts, WSL reveals the potential of weights as a learnable, structured domain and enables advances in diverse applications—including model retrieval, continual learning, federated learning, neural architecture search, and data-free reconstruction. The authors further support community progress by releasing an open-source repository dedicated to weight space research.
To address the strong data dependency, poor controllability, and limited generalization in parametric 3D shape generation, this paper proposes a data-efficient, controllable, and interpretable generative framework. Methodologically, it performs linear affine mixing in an aligned neural weight space, integrates SDF decoder overfitting with parameter-constrained optimization, and introduces a linear mismatch safety metric to ensure geometric validity. The approach achieves high-fidelity interpolation and safe extrapolation using only ~100 training samples—enabling, for the first time, full-range (100%) parameter-space extrapolation. Experiments on DrivAerNet++ and BlendedNet demonstrate substantial improvements over conditional autoencoders and Deep Neural Interpolation (DNI), particularly in data efficiency, precise parametric control, and physics-informed performance optimization.
This work investigates model collapse in overparameterized linear regression arising from iterative mixing of ground-truth labels with model-generated synthetic labels. We propose an iterative learning framework incorporating a tunable label-mixing ratio and conduct rigorous asymptotic analysis of the generalization error for both minimum-ℓ₂-norm interpolation and ridge regression. Our key theoretical finding is that, under minimum-ℓ₂-norm interpolation, the optimal proportion of ground-truth labels converges to the inverse golden ratio (≈0.618), a phenomenon governed by the geometric structure of the data covariance spectrum; this result formally establishes that model collapse is avoidable. We further derive exact asymptotic generalization error expressions, proving that the optimal mixing ratio is always at least 1/2—i.e., ground-truth data must dominate. Extensive simulations corroborate the theory, providing an interpretable, principled foundation for robust self-training.
This study addresses the prohibitive computational cost of optimizing large language model merging coefficients, which typically relies on expensive benchmark evaluations. We propose a quadratic surrogate modeling approach grounded in mixture design methodology. Specifically, classical mixture designs are employed over the coefficient simplex to determine a minimal set of measurement points, replacing random sampling with structured evaluation. A quadratic surrogate model is then constructed and interpolated over the simplex to predict multi-expert mixture performance while minimizing evaluation overhead. This method accurately predicts outcomes for unseen coefficient combinations, achieving competitive model merging performance under limited computational budgets and significantly reducing the overall cost of merging optimization.
Mixup’s blind interpolation often yields synthetic samples deviating from the underlying data manifold, degrading model calibration. We observe that interpolation distance correlates positively with mislabeling risk. To address this, we propose a similarity-driven adaptive Mixup framework that dynamically modulates the Beta distribution parameter based on feature-space distances, prioritizing interpolation between nearby samples. Our method preserves augmentation diversity while significantly mitigating manifold mismatch. Extensive experiments across multi-class classification and regression tasks demonstrate an average 35% reduction in Expected Calibration Error (ECE), consistent improvements in Brier Score, enhanced accuracy, and 2.1× faster training convergence. This work establishes the first coupling of sample similarity modeling with interpolation coefficient distribution control, introducing a novel paradigm for calibration-aware data augmentation.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
This work addresses the lack of a unified theoretical foundation in existing model merging approaches and the opacity of hyperparameters in open-source fine-tuned models, which together hinder the predictability of merged model performance. Leveraging L2-stability theory, the study establishes the first unified generalization framework to systematically analyze the generalization capability of merged heterogeneous expert models and proposes actionable fine-tuning strategies to enhance mergeability. Through parameter-space merging, derivation of generalization bounds, and large-scale vision experiments on ResNet and ViT architectures, the authors validate the critical influence of hyperparameters on merging performance across 20 and 8 tasks, respectively. Theoretical predictions align closely with empirical results, significantly improving the predictability and effectiveness of model merging.
Existing parameter reparameterization methods are often confined to a single objective—either parameter-efficient fine-tuning or model compression—making it challenging to simultaneously address both demands under resource constraints. This work proposes CRISP, a unified framework that jointly achieves model compression and parameter-efficient fine-tuning within a single architecture. CRISP decomposes pre-trained weights into shared base matrices and lightweight mixture coefficients, enhanced by cross-layer base sharing and an interpolation-based gated coefficient recombination mechanism. Requiring fewer than 200 trainable parameters, CRISP outperforms existing approaches by 1% in joint compression and fine-tuning tasks, surpasses state-of-the-art methods by up to 1.5% in pure parameter-efficient fine-tuning, and achieves a consistent 4–5% improvement in overall dual-task performance.
Existing generative models struggle to efficiently fuse heterogeneous conditions—such as high-dimensional image features and scalar-valued time/noise levels—due to architectural coupling or redundant encoding in mainstream approaches. To address this, we propose Deep Parameter Interpolation (DPI), a lightweight, plug-and-play conditioning mechanism that dynamically interpolates between two pretrained parameter sets within the neural network’s layer-wise parameter space, guided by a learnable scalar encoding. DPI is architecture-agnostic: it requires no modification to the backbone network and seamlessly integrates with diverse generative paradigms, including diffusion models and flow matching. Experiments demonstrate that DPI improves denoising accuracy and generation quality while maintaining computational overhead comparable to baseline methods. Its generalizability is validated across multiple generative tasks, confirming robust performance without sacrificing efficiency or flexibility.
This work investigates model size interpolation under zero-shot settings via layer patching, aiming to construct intermediate-scale models whose performance lies between that of a teacher and a student model. The authors formulate layer patching for the first time as a shortest-path optimization problem on a directed acyclic graph, revealing the critical impact of patching order on interpolation efficacy. They propose KLPatch, a greedy algorithm guided by KL divergence, to determine an effective patching sequence. Experimental results demonstrate that even a simple head-to-tail patching strategy yields competitive performance, while KLPatch further enhances interpolation quality. This study provides both theoretical insights and a practical method for efficiently constructing near-optimal intermediate models without additional training.
This work addresses the lack of provable generalization guarantees for multidimensional hyperparameter tuning in complex, non-smooth spaces. It proposes the first data-driven framework for such settings, establishing generalization bounds under mild assumptions by integrating structured loss with validation loss. Leveraging tools from real algebraic geometry, the analysis characterizes the complexity of semi-algebraic function classes, yielding tighter and more broadly applicable generalization bounds. The framework’s effectiveness and learnability are demonstrated on models such as weighted group Lasso and weighted fused Lasso, offering both theoretical foundations and practical methodologies for multidimensional hyperparameter optimization.