Score
Designs and implements regularizers, penalty terms, constraints, and optimization procedures that induce and control sparsity in model parameters, activations, or learned representations (e.g., L1/L0-style penalties, group or structured sparsity, thresholding/pruning schedules, and proximal updates). Builds training and distillation workflows and analyzes trade-offs among sparsity, predictive performance, and computational or memory efficiency, including methods to preserve or increase sparsity during training, across inputs, or at inference/retrieval time.
This work addresses activation redundancy in large language model (LLM) inference. We propose a post-training N:M structured activation sparsification method that integrates a lightweight error compensation mechanism, dynamic input-adaptive pruning, and multi-criterion calibration to achieve efficient I/O compression and hardware-friendly deployment. Experimental results demonstrate that the 16:32 sparsity pattern closely matches unstructured sparsity in accuracy; the 8:16 pattern reduces memory access overhead significantly while incurring less than 0.5% accuracy degradation—making it an optimal trade-off for hardware acceleration. Crucially, activation sparsification preserves generative capability more effectively than weight pruning. The method is model-agnostic and supports plug-and-play integration across diverse LLMs. Our implementation is publicly available.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
The rule of thumb regarding the relationship between the bias-variance tradeoff and model size plays a key role in classical machine learning, but is now well-known to break down in the overparameterized setting as per the double descent curve. In particular, minimum-norm interpolating estimators can perform well, suggesting the need for new tradeoff in these settings. Accordingly, we propose a regularization-sharpness tradeoff for overparameterized linear regression with an $\ell^p$ penalty. Inspired by the interpolating information criterion, our framework decomposes the selection penalty into a regularization term (quantifying the alignment of the regularizer and the interpolator) and a geometric sharpness term on the interpolating manifold (quantifying the effect of local perturbations), yielding a tradeoff analogous to bias-variance. Building on prior analyses that established this information criterion for ridge regularizers, this work first provides a general expression of the interpolating information criterion for $\ell^p$ regularizers where $p \ge 2$. Subsequently, we extend this to the LASSO interpolator with $\ell^1$ regularizer, which induces stronger sparsity. Empirical results on real-world datasets with random Fourier features and polynomials validate our theory, demonstrating how the tradeoff terms can distinguish performant linear interpolators from weaker ones.
Mixed-integer linear programming (MILP)-based formal verification of neural networks suffers from computational intractability due to inherent nonlinearity and uncertainty in embedded models. Method: This paper proposes a sparse surrogate modeling approach based on neural network pruning: the original network is pruned into a low-precision, highly sparse subnetwork—termed a “surrogate-of-a-surrogate”—which is directly encoded into the MILP framework for adversarial verification. Contribution/Results: We establish, for the first time, that pruned networks with degraded classification accuracy are inherently more amenable to MILP optimization, significantly accelerating adversarial perturbation search. Crucially, no fine-tuning is required. Experiments demonstrate substantial speedups in verification runtime while preserving soundness and completeness. This work introduces an efficient, interpretable paradigm for constrained learning and formal verification of neural networks.
The computational and energy overhead of deep learning inference is increasingly prohibitive, yet sparsity—a key optimization avenue—remains underutilized in production systems. Method: Targeting performance engineers, this work systematically surveys structured and unstructured sparsity exploitable in DNN inference and proposes an end-to-end engineering methodology—from sparse model representation to efficient sparse kernels (SpMM/SDDMM). We implement and benchmark multiple sparse computation schemes on CPU and GPU platforms, integrating support for mainstream frameworks, toolchains, and datasets. Contribution/Results: We present the first production-grade sparse inference reference framework encompassing hardware adaptation, kernel optimization, and deployment validation. Experiments demonstrate 2–5× inference speedup and substantial energy efficiency gains across representative models, establishing a reproducible, scalable practical paradigm for industrial deployment of sparse deep learning.
Existing methods often rely on soft penalties to approximate sample-level constraints, which struggle to strictly enforce hard requirements. This work proposes the first sample-wise constrained learning framework based on the sequential penalty method, enabling strict satisfaction of per-sample constraints within deep learning while providing convergence guarantees. By systematically integrating sequential penalty mechanisms into end-to-end training, the approach balances theoretical rigor with practical feasibility. Experiments on image processing tasks demonstrate that the proposed framework not only ensures strict adherence to constraints but also maintains competitive model performance.
This work addresses the computational and memory efficiency bottlenecks of large language models caused by their massive parameter counts, which often lead to significant performance degradation in existing post-training pruning methods due to the inherent lack of sparsity in the original models. To enhance pruning compatibility, the authors propose inducing sparsity at both the weight distribution and feature representation levels prior to pruning. Specifically, they introduce an absorbable equivalent scaling transformation to promote weight distribution sparsity and design a spectral norm-based loss to encourage low-rank feature sparsity. The proposed approach incurs no additional parameters or inference overhead and consistently outperforms current post-training pruning techniques across diverse model architectures and tasks, achieving highly accurate and efficient model compression.
This work investigates how large language models adapt their internal representations in response to increasing task difficulty under out-of-distribution (OOD) inputs. The study reveals that as OOD shifts intensify—manifested through more complex reasoning, longer contexts, or a greater number of answer choices—the representations in the model’s final hidden layer become markedly sparser. Building on this observation, the authors propose Sparsity-Guided In-Context Learning (SG-ICL), a curriculum-based strategy that dynamically schedules demonstration examples according to the sparsity of their representations. Extensive experiments across diverse models and tasks demonstrate a consistent correlation between representation sparsity and OOD difficulty, and show that SG-ICL effectively enhances model performance in challenging OOD scenarios.
This work addresses the lack of provable generalization guarantees for multidimensional hyperparameter tuning in complex, non-smooth spaces. It proposes the first data-driven framework for such settings, establishing generalization bounds under mild assumptions by integrating structured loss with validation loss. Leveraging tools from real algebraic geometry, the analysis characterizes the complexity of semi-algebraic function classes, yielding tighter and more broadly applicable generalization bounds. The framework’s effectiveness and learnability are demonstrated on models such as weighted group Lasso and weighted fused Lasso, offering both theoretical foundations and practical methodologies for multidimensional hyperparameter optimization.
This work addresses the challenge of unifying diverse structured sparsity patterns for efficient model compression and acceleration. The authors propose S³, an algebraic framework that formally integrates three core components—View (tensor reshaping), Block (atomic pruning units), and Scope (sparsity decision range)—to express a wide spectrum of sparsity patterns, ranging from fine-grained N:M sparsity to coarse-grained channel pruning, within a single formalism. Notably, S³ enables cross-tensor collaborative sparsification. Building upon this framework, the authors incorporate Optimal Brain Damage and Surgeon algorithms to develop structured variants of OBS/OBD. These methods significantly outperform current state-of-the-art second-order heuristic approaches in terms of output reconstruction accuracy.