Score
Designs, implements, and evaluates methods and toolchains to reduce the storage, memory footprint, and computational cost of machine learning models while preserving target accuracy or other performance metrics. This includes techniques and analyses for pruning, quantization, knowledge distillation, low-rank factorization, weight sharing, compression-aware training, and hardware-aware optimization to trade off model size, latency, energy, and accuracy.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
To address the trade-off among model size, inference latency, and accuracy degradation when deploying deep neural networks on edge devices, this paper proposes two co-designed pruning-quantization joint optimization frameworks. Methodologically, it tightly integrates feature-map similarity–based filter pruning with adaptive power-of-two (APoT) quantization, jointly optimizing pruning masks and low-bit (≤4-bit) quantization parameters during training. The key contribution lies in leveraging the complementarity of pruning and APoT: pruning eliminates structural redundancy, while APoT enhances quantized representation efficiency—thereby avoiding error accumulation inherent in sequential compression. Experiments on ResNet and VGG demonstrate that our approach achieves 5.2× model size reduction, 6.8× FLOPs reduction, and 4.1× inference speedup, with ≤0.3% Top-1 accuracy drop relative to full-precision baselines—substantially outperforming standalone pruning or quantization methods and exhibiting strong practicality for edge deployment.
Large pre-trained models incur substantial energy consumption and environmental impact, yet existing model compression research primarily prioritizes accuracy preservation without quantifying real-world electricity usage. This work establishes, for the first time, a direct empirical link between model compression techniques and measured power consumption. We systematically evaluate three classes of structural compression—pruning, low-rank decomposition, and steganographic capacity reduction—across nine pre-trained models (8M–138M parameters) under standardized training conditions. Results show that steganographic capacity reduction achieves an average 37% reduction in training energy consumption with <0.8% accuracy degradation, whereas conventional pruning and low-rank decomposition yield negligible energy savings. We introduce a reproducible, hardware-level power monitoring experimental framework and uncover a nonlinear relationship between structural compression pathways and energy efficiency. This work provides a foundational methodology and empirical evidence for green AI, shifting the evaluation paradigm from pure accuracy to energy-aware model design.
This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.
The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.
This work addresses the discrepancy between conventional compression metrics—such as parameter count and FLOPs—and actual inference latency in CPU- and memory-constrained edge deployment scenarios, where such proxies often fail to reflect real-world performance. To bridge this gap, the authors propose a latency-driven, sequential compression pipeline that integrates unstructured pruning, INT8 quantization-aware training (QAT), and knowledge distillation (KD) within a unified training framework, jointly optimizing model accuracy, size, and inference speed. Experimental results demonstrate that this specific ordering significantly outperforms alternative combinations, achieving CPU inference latencies of 0.99–1.42 milliseconds on CIFAR-10/100 with ResNet-18, WRN-28-10, and VGG-16-BN models while maintaining high accuracy and compactness, thereby establishing a new paradigm for edge-oriented model compression under realistic latency constraints.
Machine learning systems exhibit significant software-layer energy inefficiency due to redundant or suboptimal operator implementations, yet effective diagnostic tools remain lacking. This paper introduces differential energy debugging—a novel methodology that establishes the first operator-level differential energy analysis framework. It automatically identifies high-energy code regions and configuration flaws by comparing energy consumption across structurally equivalent models deployed on different ML frameworks. The approach integrates fine-grained energy profiling, cross-framework energy normalization, and root-cause inference. Evaluated on mainstream frameworks including PyTorch and TensorFlow, it detects 24 energy inefficiencies across nine widely used ML systems—including eight previously unknown defects. Seven of these were confirmed and fixed by developers, uncovering long-overlooked software-level energy bottlenecks. Our work provides a practical, actionable debugging paradigm for green AI development.
Complex models in healthcare applications suffer from high computational latency and resource consumption during inference. Method: This paper proposes an adaptive quantization strategy tailored to medical data, systematically investigating the impact of numerical precision reduction (from float64 to float32 or int32) on logistic regression performance. It characterizes the trade-off between precision compression and predictive accuracy, identifying key parameter dependencies while preserving model architecture and ensuring hardware compatibility and clinical reliability. Contribution/Results: Experiments across multiple real-world medical datasets demonstrate a 40–65% reduction in inference latency with only a 0.3–1.2% decrease in AUC, confirming both efficiency and robustness. The approach provides a reproducible, practical pathway for deploying lightweight, trustworthy AI models in resource-constrained clinical environments.