model compression

Designs, implements, and evaluates methods and toolchains to reduce the storage, memory footprint, and computational cost of machine learning models while preserving target accuracy or other performance metrics. This includes techniques and analyses for pruning, quantization, knowledge distillation, low-rank factorization, weight sharing, compression-aware training, and hardware-aware optimization to trade off model size, latency, energy, and accuracy.

modelcompression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

To address the trade-off among model size, inference latency, and accuracy degradation when deploying deep neural networks on edge devices, this paper proposes two co-designed pruning-quantization joint optimization frameworks. Methodologically, it tightly integrates feature-map similarity–based filter pruning with adaptive power-of-two (APoT) quantization, jointly optimizing pruning masks and low-bit (≤4-bit) quantization parameters during training. The key contribution lies in leveraging the complementarity of pruning and APoT: pruning eliminates structural redundancy, while APoT enhances quantized representation efficiency—thereby avoiding error accumulation inherent in sequential compression. Experiments on ResNet and VGG demonstrate that our approach achieves 5.2× model size reduction, 6.8× FLOPs reduction, and 4.1× inference speedup, with ≤0.3% Top-1 accuracy drop relative to full-precision baselines—substantially outperforming standalone pruning or quantization methods and exhibiting strong practicality for edge deployment.

Addressing computational and memory demands on resource-constrained devicesCombining pruning and quantization for efficient DNN compressionPreserving model accuracy while achieving higher compression efficiency

Energy Considerations for Large Pretrained Neural Networks

Jun 02, 2025
LM
Leo Mei
🏛️ San Jose State University

Large pre-trained models incur substantial energy consumption and environmental impact, yet existing model compression research primarily prioritizes accuracy preservation without quantifying real-world electricity usage. This work establishes, for the first time, a direct empirical link between model compression techniques and measured power consumption. We systematically evaluate three classes of structural compression—pruning, low-rank decomposition, and steganographic capacity reduction—across nine pre-trained models (8M–138M parameters) under standardized training conditions. Results show that steganographic capacity reduction achieves an average 37% reduction in training energy consumption with <0.8% accuracy degradation, whereas conventional pruning and low-rank decomposition yield negligible energy savings. We introduce a reproducible, hardware-level power monitoring experimental framework and uncover a nonlinear relationship between structural compression pathways and energy efficiency. This work provides a foundational methodology and empirical evidence for green AI, shifting the evaluation paradigm from pure accuracy to energy-aware model design.

Comparing energy usage of uncompressed and compressed modelsEvaluating compression techniques for energy efficiencyReducing electricity consumption in large neural networks

Latest Papers

What's happening recently
View more

This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.

computational costenergy consumptionenvironmental sustainability

The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.

data movementenergy efficiencyhardware-software co-design

This work addresses the discrepancy between conventional compression metrics—such as parameter count and FLOPs—and actual inference latency in CPU- and memory-constrained edge deployment scenarios, where such proxies often fail to reflect real-world performance. To bridge this gap, the authors propose a latency-driven, sequential compression pipeline that integrates unstructured pruning, INT8 quantization-aware training (QAT), and knowledge distillation (KD) within a unified training framework, jointly optimizing model accuracy, size, and inference speed. Experimental results demonstrate that this specific ordering significantly outperforms alternative combinations, achieving CPU inference latencies of 0.99–1.42 milliseconds on CIFAR-10/100 with ResNet-18, WRN-28-10, and VGG-16-BN models while maintaining high accuracy and compactness, thereby establishing a new paradigm for edge-oriented model compression under realistic latency constraints.

inference latencyknowledge distillationneural network compression

Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging

Dec 09, 2025
YP
Yi Pan
🏛️ University of Washington | Boston University | Shanghai Jiao Tong University

Machine learning systems exhibit significant software-layer energy inefficiency due to redundant or suboptimal operator implementations, yet effective diagnostic tools remain lacking. This paper introduces differential energy debugging—a novel methodology that establishes the first operator-level differential energy analysis framework. It automatically identifies high-energy code regions and configuration flaws by comparing energy consumption across structurally equivalent models deployed on different ML frameworks. The approach integrates fine-grained energy profiling, cross-framework energy normalization, and root-cause inference. Evaluated on mainstream frameworks including PyTorch and TensorFlow, it detects 24 energy inefficiencies across nine widely used ML systems—including eight previously unknown defects. Seven of these were confirmed and fixed by developers, uncovering long-overlooked software-level energy bottlenecks. Our work provides a practical, actionable debugging paradigm for green AI development.

Detects and diagnoses excessive energy use at operator levelIdentifies software energy waste from poor design in ML frameworksOptimizes energy efficiency in ML systems via differential debugging

Complex models in healthcare applications suffer from high computational latency and resource consumption during inference. Method: This paper proposes an adaptive quantization strategy tailored to medical data, systematically investigating the impact of numerical precision reduction (from float64 to float32 or int32) on logistic regression performance. It characterizes the trade-off between precision compression and predictive accuracy, identifying key parameter dependencies while preserving model architecture and ensuring hardware compatibility and clinical reliability. Contribution/Results: Experiments across multiple real-world medical datasets demonstrate a 40–65% reduction in inference latency with only a 0.3–1.2% decrease in AUC, confirming both efficiency and robustness. The approach provides a reproducible, practical pathway for deploying lightweight, trustworthy AI models in resource-constrained clinical environments.

Applying optimization methods to healthcare datasets for efficiencyOptimizing machine learning models through quantization techniquesReducing time complexity while preserving model accuracy

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
MM

Michael Milford

QUT Professor | Director, QUT Robotics Centre | ARC Laureate Fellow | Microsoft Fellow
Roboticscomputational neurosciencenavigationSLAM
JW

Jiacheng Wang

Nanyang Technological University
ISACGenAILow-altitude wireless networkSemantic Communications
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management