Score
Designs and implements algorithms and procedures to compute the importance or similarity of model components and selectively remove less-important or redundant parts of a machine learning model while preserving specified task knowledge. This includes building efficient importance estimators (for example using a single forward/backward pass), selection and pruning strategies, and analyses that measure retained-task performance and the rate of forgetting convergence.
This work addresses the challenge of deploying multi-component neural network controllers, which are often hindered by high computational complexity, and the inadequacy of conventional norm-based pruning methods in accurately capturing the functional importance of individual components. To this end, the paper introduces a component-aware structured pruning framework that, for the first time, integrates three gradient-driven importance metrics—gradient accumulation, Fisher information, and Bayesian uncertainty—into the pruning of multi-component controllers. These metrics enable dynamic assessment of component importance during training, uncovering structural dependencies and temporal variations overlooked by static heuristic approaches. Experiments on autoencoders and TD-MPC reinforcement learning agents demonstrate that the proposed method more accurately identifies critical components, achieving substantial model compression while effectively preserving performance.
This work proposes a joint optimization pruning strategy that moves beyond the conventional paradigm of assessing weight importance solely based on magnitude. Recognizing that performance degradation caused by weight removal can be partially compensated through adjustments to adjacent biases, the method simultaneously prunes weights and computes optimal bias perturbations via automatic differentiation to minimize accuracy loss. By explicitly accounting for the interplay between weights and biases, this approach provides a more accurate measure of weight significance. Extensive experiments demonstrate that the proposed technique consistently outperforms state-of-the-art pruning methods across various models and tasks, achieving superior accuracy and robustness—particularly under high pruning ratios.
This study addresses the limitation of relying solely on expert utilization rates to assess removal damage during expert pruning in Mixture-of-Experts (MoE) models, emphasizing the necessity of preserving functional substitutability to maintain output distributions. To this end, it proposes a training-free expert pruning framework that introduces a consensus residual-based scoring mechanism for evaluating functional substitutability. By integrating an exact single-deletion identity with calibration token aggregation, the method achieves precise pruning without requiring gradients or recovery training. Extensive experiments across multiple large-scale models and varying pruning ratios demonstrate that the proposed approach attains the highest macro-average score over nine evaluation tasks. It significantly outperforms the REAP baseline while effectively reducing reverse KL divergence, highlighting its efficacy in maintaining model performance under aggressive pruning conditions.
Existing data pruning methods suffer from complex designs and poorly understood mechanisms of their key components, hindering progress in the field. Method: This paper pioneers a decoupled framework that separates data pruning into two orthogonal modules—“data representation” and “selection algorithm”—and systematically evaluates their independent contributions to instance selection efficacy in NLP model training. Through theoretical analysis and extensive empirical evaluation across multiple benchmarks (including gradient- and embedding-based representations and diverse selection algorithms), we assess their relative impact. Contribution/Results: We find that representation quality dominates selection algorithm choice: high-fidelity representations (e.g., training gradients) substantially improve pruning effectiveness, whereas no selection algorithm exhibits consistent superiority across tasks; even for identical objectives, different algorithms yield markedly divergent selected subsets. This work establishes an interpretable, reusable analytical framework for data pruning and identifies representation optimization—not algorithmic refinement—as the primary lever for enhancing pruning efficiency.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
研究解决了过分散路由下专家修剪导致模型性能下降的问题,提出MESA方法以最小化最坏情况下的领域退化。
This study investigates whether neurons in task-specific large language models contribute uniformly to target tasks and proposes an efficient pruning method to reduce computational overhead. Through systematic neuron ablation experiments, the authors employ activation-based selective metrics to identify low-contribution neurons and compare this approach against random pruning. The work provides the first empirical evidence that critical task information is concentrated in a small subset of neurons: removing approximately 10% of these key neurons leads to catastrophic performance degradation, whereas selective pruning can eliminate 30%–35% of parameters while preserving performance. Further fine-tuning effectively recovers task accuracy, substantially reducing model size, GPU memory consumption, and significantly improving inference throughput.
研究通过动态专家剪枝方法在细粒度MoE架构中减少冗余专家选择,保留约2/3专家即可保持98.8%性能,提高推理效率。
本文通过定义和分析不同类型的零重要性,解决了解释机器学习中特征相关性的多种概念问题,从而为特征分析提供了一个统一的框架。
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.