Score
Designs and implements compact models and the training/fine-tuning pipelines that specialize them directly on-device or at the network edge under strict compute, memory, latency, and energy constraints. Builds and analyzes resource-aware methods—such as architecture search, pruning, quantization, distillation, and lightweight online adaptation—that enable small models (e.g., ~10k parameters) to adapt on‑the‑fly to current context and match or exceed larger generalist performance.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
This work addresses the challenges of deploying Transformer models on resource-constrained edge devices, where computational complexity, memory footprint, and power consumption pose significant bottlenecks. The study systematically evaluates lightweight Transformer architectures alongside optimization strategies—including compression, quantization, pruning, and knowledge distillation—and integrates sparse attention mechanisms, mixed-precision quantization (INT8/FP16), and hardware-aware neural architecture search to enable efficient deployment within frameworks such as TensorFlow Lite and CoreML. A proposed six-step deployment pipeline achieves 4–10× model compression and 3–9× latency reduction at a power budget of 2–5 W, with accuracy degradation below 2% (retaining 75–96% of original accuracy). The analysis further uncovers a consistent memory bandwidth bottleneck, revealing that models with 15–40 million parameters attain 60–75% hardware utilization on mainstream edge platforms.
Edge-device GPUs face severe memory constraints, hindering fine-tuning and multimodal extension of large language models (LLMs). Method: We systematically survey memory-efficient fine-tuning techniques (e.g., LoRA, QLoRA, Adapters) and model compression methods (e.g., quantization, pruning, knowledge distillation, sparse training), and propose, for the first time, a synergistic fine-tuning-and-compression paradigm tailored for edge deployment. We design a unified evaluation framework that quantifies trade-offs across three dimensions: energy efficiency, hardware compatibility, and multimodal generalization capability. Contribution/Results: We establish the first taxonomy of LLM lightweighting techniques specifically for edge deployment, characterizing each method’s performance in GPU memory footprint, inference latency, accuracy retention, and cross-platform adaptability. Our work provides both theoretical foundations and practical guidelines for sustainable on-device AI deployment.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
To address the challenge of fine-tuning large language models (LLMs) on memory- and compute-constrained edge devices, this paper proposes an efficient on-device fine-tuning method that requires no modifications to the inference engine. Our approach comprises three key contributions: (1) Parallelized Random Gradient Estimation (P-RGE), a low-overhead gradient approximation technique operating within a zeroth-order optimization framework; (2) a lightweight LoRA-FA module fully compatible with the ExecuTorch runtime, requiring no intrusive changes to the execution stack; and (3) the synergistic integration of LoRA-based parameter-efficient fine-tuning with P-RGE, achieving up to 68% reduction in GPU memory consumption and significantly lower computational overhead. Experiments demonstrate that our method maintains fine-tuning accuracy while accelerating training by 3.2×, enabling real-time, personalized LLM deployment on edge devices. This work provides a practical pathway for continual learning of LLMs in resource-constrained environments.
Large-scale collaborative model training across resource-constrained, heterogeneous edge devices faces challenges including data decentralization, computational heterogeneity, knowledge loss due to unstructured pruning, and straggler-induced computation bottlenecks. To address these, we propose Co-S²P, a semi-asynchronous collaborative training framework that innovatively integrates data-distribution-aware structured pruning with cross-module knowledge distillation. Co-S²P enables resource-adaptive submodel generation and semi-asynchronous parameter updates, and provides an asymptotically optimal convergence rate of O(1/√(N·E·Q)), where N, E, and Q denote the number of devices, local epochs, and pruning granularity, respectively. Evaluated on 16 NVIDIA Jetson devices, Co-S²P achieves up to 8.8% higher accuracy, 22% lower memory footprint, 24% faster training time, and 1.2× improved resource utilization compared to state-of-the-art baselines.
This study addresses the mismatch between theoretical compression efficacy and empirical performance in deploying large language models on edge devices, alongside the absence of practical deployment guidelines. Through extensive multi-hardware benchmarking integrating quantization, pruning, LoRA, and latency decomposition, we systematically evaluate diverse compression strategies. Our analysis reveals complex trade-offs between latency and accuracy, demonstrating that structured pruning can paradoxically increase latency by 3.4× while inducing spurious capabilities. Consequently, we distill task-specific, actionable guidelines for efficient edge deployment and release our codebase as open source. This work provides critical theoretical insights and practical references to facilitate the real-world implementation of edge AI, bridging the gap between model compression research and hardware-aware deployment optimization.
This work addresses the challenge of efficiently constructing and managing massive numbers of persistent personalized models atop trillion-parameter foundation models. It proposes leveraging parameter-efficient fine-tuning (PEFT) as a lightweight and reliable personalization substrate, combining a shared large model with small, trainable adapters to encode user preferences, skills, and memory. The authors introduce MinT, an infrastructure that integrates adapter identity management, version control, provenance tracking, evaluation, and serving mechanisms, and define three scaling dimensions: Scale Up, Scale Down, and Scale Out. Experimental results demonstrate that, even under strong shared priors, compact adapters can stably capture personalized behaviors, offering a viable pathway toward large-scale deployment of millions of persistent personal models.
Despite growing interest in Mixture-of-Experts (MoE) models, large-scale end-to-end pretraining of MoE architectures on pure AMD hardware—specifically the MI300X GPU and Pollara interconnect—remains unexplored, raising questions about the platform’s readiness for state-of-the-art LLM training. Method: We introduce MI300X-aware Transformer module sizing guidelines, conduct full-stack Pollara communication microbenchmarks, and integrate optimized All-reduce/Reduce-scatter primitives, fault-tolerant training, and checkpoint reshaping. Contribution/Results: We successfully complete end-to-end pretraining of the ZAYA1-base MoE model (760M activated parameters, 8.3B total parameters) on native AMD infrastructure. Evaluated on reasoning, mathematics, and code generation, ZAYA1-base substantially outperforms Llama-3-8B and OLMoE, matching the performance of Qwen3-4B and Gemma3-12B. This demonstrates, for the first time, that AMD’s hardware-software stack achieves full maturity for high-performance MoE model training.
This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.
CIM macros suffer from low throughput and high inference error due to physical area constraints and limited ADC precision. To address this, we propose a two-stage model adaptation framework: (1) layer-importance-aware model compression and resource reallocation to maximize CIM array utilization; and (2) quantization-aware training integrated with partial-sum quantization modeling to explicitly compensate for ADC non-idealities. Our approach is the first to enable layer-importance-driven co-optimization of CIM resources and supports concurrent activation of 256 wordlines. Experiments demonstrate a 93% model compression ratio, 90% array utilization, inference accuracy on par with floating-point baselines, and significantly reduced weight loading latency.