Score
Designs, builds, and analyzes inference systems, runtimes, and models that minimize end-to-end latency—particularly batch-1 execution—on-device or at the edge by applying operator fusion, low-overhead execution strategies, and memory/execution-footprint reductions. Work includes measuring and profiling inference latency, implementing fused and batch-1-optimized operators, and optimizing deployment for embedded/edge hardware to meet strict latency constraints.
Deploying complex AI models efficiently on resource-constrained edge devices (e.g., smartphones, IoT endpoints) remains challenging due to stringent latency, memory, and energy constraints. Method: This paper proposes a three-tier co-optimization paradigm—spanning data, model, and system layers—unifying cross-layer coordination mechanisms for the first time and bridging a critical methodological gap in end-to-cloud full-stack optimization. It integrates lightweight edge data preprocessing, model compression techniques (pruning, quantization, knowledge distillation), efficient inference runtimes (TensorFlow Lite, ONNX Runtime), and heterogeneous hardware acceleration (NPU/GPU/FPGA). Contribution/Results: We construct a reusable edge AI optimization roadmap, achieving an average 3.2× inference speedup and 78% memory reduction while preserving ≥95% of the original model accuracy.
To address resource constraints and hard real-time requirements in multi-mobile-device DNN inference offloading to GPU-enabled edge servers, this paper proposes a joint optimization framework that simultaneously determines task offloading decisions, DNN-layer-granularity subtask partitioning, GPU dynamic batching, and DVFS-based frequency scaling. We formulate the problem as a mixed-integer nonlinear program (MINLP) — the first such model for this setting — and design the low-complexity J-DOB algorithm, which provides theoretical guarantees on near-optimal performance. Compared to purely local execution, J-DOB reduces total system energy consumption by 51.30% and 45.27% under identical and heterogeneous deadline settings, respectively, while significantly improving both energy efficiency and end-to-end latency compliance. The core contribution lies in the four-dimensional coupled modeling and efficient optimization of layer partitioning, batch scheduling, DVFS, and offloading decisions.
This work addresses the challenge of deploying large language models on edge devices, which is hindered by a lack of systematic understanding of inference latency and energy efficiency scaling across heterogeneous hardware (CPU/GPU/NPU). The authors propose QEIL, a unified framework that, for the first time, uncovers stable power-law scaling behaviors of Transformer models with respect to latency, energy consumption, and task coverage. Leveraging these insights, QEIL introduces three composite metrics and a safety-aware intelligent scheduler to enable coordinated optimization across heterogeneous accelerators from diverse vendors. Through formal modeling, computational orchestration, thermal management, fault-tolerant execution, and hardware health monitoring, QEIL significantly improves energy efficiency, reduces latency, and expands task coverage across five model families—while preserving model accuracy and ensuring system safety.
Operator composition and performance trade-offs remain open challenges in multi-tier edge AI deployment. Method: This paper systematically evaluates the black-box deployment efficacy of three operator classes—model partitioning, quantization, and early-exit—across mobile-edge-cloud architectures, conducting cross-device empirical studies on ResNet, ResNeXt, DUC, and FCN vision models. Contribution/Results: We first reveal that a quantization + early-exit hybrid strategy achieves optimal latency–accuracy trade-off at the edge with only marginal accuracy degradation (<2%). We propose a scenario-aware deployment decision framework grounded in resource constraints, network conditions, and input scale: (i) mobile–edge collaborative partitioning under tight resource budgets; (ii) cloud-only execution for small-scale inputs; and (iii) edge-first execution for large-scale inputs. Experimental results demonstrate that quantization alone better preserves accuracy, whereas hybrid strategies significantly improve edge energy efficiency.
To address the insufficient inference capability of resource-constrained edge devices, this paper systematically compares hierarchical inference (HI) against pure on-device inference across accuracy, latency, and energy consumption. Existing HI studies overlook device-side latency and energy overhead and fail to jointly model the heterogeneous hardware, network, and model dimensions. To bridge this gap, we conduct the first multi-dimensional empirical evaluation of HI on real embedded hardware. We propose Early Exit with HI (EE-HI), a hybrid mechanism that dynamically coordinates lightweight local models and remote servers: it adaptively offloads samples and enables early exit during image classification, thereby eliminating HI’s inherent fixed overhead. Experiments show that HI reduces latency by up to 73% and device energy consumption by up to 77% compared to pure on-device inference; EE-HI further improves upon HI by reducing latency by 59.7% and device energy consumption by 60.4%.
Existing edge AI inference systems are constrained by model-level mapping strategies, which hinder efficient utilization of heterogeneous computing resources to accommodate diverse operator characteristics. This work proposes the first unified operator-level scheduling framework that dynamically assigns each operator to the optimal processing unit (CPU/GPU/NPU) based on empirical performance profiling. By constructing a weighted execution graph and solving a shortest-path problem, the framework enables latency- or energy-efficiency-oriented scheduling. It transcends conventional limitations by uniformly supporting sequential execution, intra-model parallelism, and multi-model concurrency, all without relying on model-specific heuristics, thus achieving model-agnostic applicability. Experiments on an Intel Core Ultra SoC demonstrate up to 1.60× speedup with intra-model parallelism, a geometric mean acceleration of 3.42× for concurrent multi-model execution, and an average energy saving of 48.2% under energy-efficient scheduling.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
This study addresses the stringent constraints on size, power, and computational resources faced by AI inference on resource-limited platforms such as small satellites. By conducting empirical characterization of quantized AI inference on Cortex-M-class processors using representative embedded vision neural networks, the work establishes the first measurement-based performance baseline for on-board embedded systems. It introduces an explicit multi-core/multi-device cooperative scheduling mechanism and integrates analysis of ALU/SIMD utilization with memory traffic to evaluate system behavior. Moving beyond conventional paradigms that rely on opaque OS-level scheduling, this research provides comparable latency and data-movement benchmarks for typical spaceborne processors like LEON and NOEL-V, thereby demonstrating the critical role of architecture-aware design and cooperative scheduling as key dimensions in optimizing embedded AI inference for satellite applications.
This work addresses the challenge of accurately predicting inference latency under dynamic voltage and frequency scaling (DVFS) on mobile edge devices, where fluctuating CPU/GPU frequencies render traditional static analysis ineffective and exhaustive empirical profiling prohibitively expensive—particularly for small language models (SLMs) with variable context lengths. To overcome this, the authors propose FLAME, a novel method that introduces the first fine-grained model of asynchronous CPU-GPU execution. FLAME enables bottom-up, frequency-aware latency prediction across both deep neural networks (DNNs) and SLMs by combining layer-level delay decomposition, quantification of parallel execution overlap and pipeline bubbles, and frequency extrapolation from sparsely sampled measurements. The approach reduces modeling time from hours or days to minutes, drastically cuts required empirical samples, maintains low prediction error, and enhances deadline-aware DVFS scheduling by jointly optimizing energy efficiency and latency guarantees, outperforming existing solutions.