Score
Designs, implements, and validates optimized on‑device inference artifacts and deployment pipelines for edge hardware, including model compression and conversion (e.g., quantization, pruning), selection and tuning of compilers/backends and runtime kernels, and creation of device-installable packages or firmware/OTA bundles. Measures and analyzes latency, memory, energy, and throughput trade-offs, and builds tooling to automate packaging, versioning, and safe rollout to constrained devices.
Deploying complex AI models efficiently on resource-constrained edge devices (e.g., smartphones, IoT endpoints) remains challenging due to stringent latency, memory, and energy constraints. Method: This paper proposes a three-tier co-optimization paradigm—spanning data, model, and system layers—unifying cross-layer coordination mechanisms for the first time and bridging a critical methodological gap in end-to-cloud full-stack optimization. It integrates lightweight edge data preprocessing, model compression techniques (pruning, quantization, knowledge distillation), efficient inference runtimes (TensorFlow Lite, ONNX Runtime), and heterogeneous hardware acceleration (NPU/GPU/FPGA). Contribution/Results: We construct a reusable edge AI optimization roadmap, achieving an average 3.2× inference speedup and 78% memory reduction while preserving ≥95% of the original model accuracy.
This paper addresses core challenges in deploying AI models on edge and end devices—namely, poor real-time performance, severe resource constraints, and data privacy risks. To this end, it proposes the first unified definition and multidimensional analytical framework for on-device AI, integrating dual perspectives from edge computing and large language model evolution, and establishing a systematic “end-to-end challenges–technical solutions” mapping. Methodologically, it comprehensively synthesizes key techniques including model compression (pruning, quantization, distillation), lightweight architecture design, hardware-software co-optimization, edge-side preprocessing, and federated learning. Contributions include: (1) categorizing 12 representative application scenarios; (2) identifying 7 fundamental technical bottlenecks; and (3) distilling 4 practical, deployable evolutionary pathways. The resulting structured survey serves as an authoritative benchmark and implementation guide for both industrial deployment and academic research.
This paper addresses the deployment of multi-model inference pipelines on resource-constrained edge devices. We propose an end-to-end adaptive configuration framework that, for the first time, explicitly incorporates device resource constraints into joint optimization decisions. Our method integrates residual feature extraction, LSTM-based workload forecasting, and policy-gradient reinforcement learning to jointly optimize QoS guarantees (e.g., latency and throughput), operational cost, and real-time adaptability. Evaluated on a real Kubernetes-based edge cluster, the framework achieves a 27% reduction in average inference latency, a 31% increase in throughput, a 22% decrease in deployment cost, and over 42% faster configuration decision-making for complex pipelines—outperforming state-of-the-art baselines. The core contributions are: (i) a resource-aware joint optimization model that unifies hardware constraints with pipeline scheduling and scaling decisions; and (ii) a lightweight, learning-driven configuration mechanism enabling efficient, online adaptation under dynamic edge conditions.
This work addresses the challenge of deploying large language models on edge devices, which is hindered by a lack of systematic understanding of inference latency and energy efficiency scaling across heterogeneous hardware (CPU/GPU/NPU). The authors propose QEIL, a unified framework that, for the first time, uncovers stable power-law scaling behaviors of Transformer models with respect to latency, energy consumption, and task coverage. Leveraging these insights, QEIL introduces three composite metrics and a safety-aware intelligent scheduler to enable coordinated optimization across heterogeneous accelerators from diverse vendors. Through formal modeling, computational orchestration, thermal management, fault-tolerant execution, and hardware health monitoring, QEIL significantly improves energy efficiency, reduces latency, and expands task coverage across five model families—while preserving model accuracy and ensuring system safety.
This work addresses the challenge of on-device model adaptation under resource constraints, where end-to-end backpropagation through modern deep networks is infeasible on edge hardware. The authors propose a heterogeneous adaptation pipeline that deploys an INT8-quantized backbone on a Hailo-8L inference accelerator for frozen feature extraction, while fine-tuning only a lightweight classification head in FP32 on the CPU. A post-training quantization recovery mechanism is introduced to preserve feature quality. This approach uniquely repurposes commercial inference accelerators for on-device training and overcomes training efficiency bottlenecks through co-design of the backbone and head. Experiments demonstrate up to a 15.4× speedup in training over a Raspberry Pi 5 CPU baseline, substantially reduced per-sample energy consumption, and competitive accuracy across multiple architectures and datasets.
Adaptive model deployment on heterogeneous edge devices remains challenging due to cumbersome model-aware design and time-consuming hardware-aware analysis in existing SuperNet approaches. Method: This paper proposes the first end-to-end automated SuperNet deployment framework: it leverages computation-graph-guided compilation to automatically transform arbitrary user models into lightweight supernets; and integrates learning-free latency and accuracy predictors for zero-shot, low-overhead cross-hardware performance estimation and model specialization. Contributions/Results: Compared to state-of-the-art methods, our framework reduces supernet code size by 11–27×, cuts hardware tuning cost by over 11×, improves absolute accuracy by up to 15.60%, and reduces inference latency by 60.03%.
This work addresses the challenges of deploying Transformer models on resource-constrained edge devices, where computational complexity, memory footprint, and power consumption pose significant bottlenecks. The study systematically evaluates lightweight Transformer architectures alongside optimization strategies—including compression, quantization, pruning, and knowledge distillation—and integrates sparse attention mechanisms, mixed-precision quantization (INT8/FP16), and hardware-aware neural architecture search to enable efficient deployment within frameworks such as TensorFlow Lite and CoreML. A proposed six-step deployment pipeline achieves 4–10× model compression and 3–9× latency reduction at a power budget of 2–5 W, with accuracy degradation below 2% (retaining 75–96% of original accuracy). The analysis further uncovers a consistent memory bandwidth bottleneck, revealing that models with 15–40 million parameters attain 60–75% hardware utilization on mainstream edge platforms.
This work addresses the lack of systematic investigation into hardware reliability under input degradation when deploying deep learning models on resource-constrained edge devices. The authors propose a decoupled fault injection framework that, for the first time, integrates large language models (LLMs) with latent diffusion models (LDMs) to generate realistic degraded inputs. Using this framework, they conduct multidimensional hardware monitoring—covering CPU/GPU utilization, memory usage, power consumption, throughput, and temperature—on a Jetson Nano platform running TensorRT-optimized YOLOv10s, YOLOv11s, and YOLO2026n models. Experimental results demonstrate that even under severely degraded inputs, TensorRT inference maintains stable GPU occupancy, controlled temperature rise, and safe power consumption, while memory usage exhibits predictable release patterns after warm-up, providing empirical evidence to guide robust edge AI system design.
This work addresses the challenge of end-to-end machine learning inference on microcontroller-class edge devices under stringent constraints on memory, energy consumption, and latency. To bridge the gap between conventional machine learning pipelines and embedded deployment realities, the authors propose a robust design framework tailored for resource-constrained environments, encompassing data acquisition, preprocessing, model compression, and streaming deployment. The framework integrates sampling buffers, feature dimensionality reduction techniques (e.g., RMS, spectral features, MFCCs), validation strategies for class imbalance, and co-optimization of models with runtime systems to form a complete embedded ML pipeline. Experimental evaluations on two representative tasks—inertial human activity recognition and keyword spotting—demonstrate that the proposed approach enables efficient, practical, and robust on-device inference, significantly narrowing the divide between general-purpose machine learning methodologies and embedded implementation requirements.
This work addresses the challenge of performing end-to-end on-device training and inference for vision-based machine learning on extremely resource-constrained microcontrollers, where traditional approaches rely heavily on cloud infrastructure. The authors demonstrate a complete pipeline—including data acquisition, training of a two-layer convolutional neural network (CNN), and real-time inference—on an ESP32-S3 XIAO ML Kit featuring only 8 MB of PSRAM and no external dependencies. By leveraging batch-level gradient accumulation, precomputed scaling lookup tables, a three-priority weight loading scheme, and PSRAM-aware memory management, the system achieves a full training cycle in just 9 minutes and inference at 6.3 frames per second on a 64×64 three-class classification task. Implemented in only 1,750 lines of C++ code, the system includes a custom Adam optimizer and CNN compatible with the Arduino IDE, and is released under the MIT license.
This work addresses the dual challenges posed by the complexity of deep learning models and the heterogeneity of edge devices, which render conventional hardware-agnostic approaches inadequate in balancing efficiency and performance. To overcome these limitations, the study proposes a hardware-aware co-design paradigm that integrates model compression with neural architecture search to tailor efficient model architectures specifically for target edge hardware. By moving beyond one-size-fits-all deployment strategies, the proposed method breaks through the performance bottlenecks of generic solutions, significantly enhancing inference efficiency and resource utilization across diverse heterogeneous edge platforms. This enables more efficient and scalable deployment of intelligent systems at the edge.