latency-first inference

Designs, builds, and analyzes inference systems, runtimes, and models that minimize end-to-end latency—particularly batch-1 execution—on-device or at the edge by applying operator fusion, low-overhead execution strategies, and memory/execution-footprint reductions. Work includes measuring and profiling inference latency, implementing fused and batch-1-optimized operators, and optimizing deployment for embedded/edge hardware to meet strict latency constraints.

latency-firstinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$224K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address resource constraints and hard real-time requirements in multi-mobile-device DNN inference offloading to GPU-enabled edge servers, this paper proposes a joint optimization framework that simultaneously determines task offloading decisions, DNN-layer-granularity subtask partitioning, GPU dynamic batching, and DVFS-based frequency scaling. We formulate the problem as a mixed-integer nonlinear program (MINLP) — the first such model for this setting — and design the low-complexity J-DOB algorithm, which provides theoretical guarantees on near-optimal performance. Compared to purely local execution, J-DOB reduces total system energy consumption by 51.30% and 45.27% under identical and heterogeneous deadline settings, respectively, while significantly improving both energy efficiency and end-to-end latency compliance. The core contribution lies in the four-dimensional coupled modeling and efficient optimization of layer partitioning, batch scheduling, DVFS, and offloading decisions.

Balance latency constraints with GPU batching and offloadingMinimize total energy consumption under strict deadline requirementsOptimize energy in edge-device co-inference for mobile DNN tasks

This work addresses the challenge of deploying large language models on edge devices, which is hindered by a lack of systematic understanding of inference latency and energy efficiency scaling across heterogeneous hardware (CPU/GPU/NPU). The authors propose QEIL, a unified framework that, for the first time, uncovers stable power-law scaling behaviors of Transformer models with respect to latency, energy consumption, and task coverage. Leveraging these insights, QEIL introduces three composite metrics and a safety-aware intelligent scheduler to enable coordinated optimization across heterogeneous accelerators from diverse vendors. Through formal modeling, computational orchestration, thermal management, fault-tolerant execution, and hardware health monitoring, QEIL significantly improves energy efficiency, reduces latency, and expands task coverage across five model families—while preserving model accuracy and ensuring system safety.

Edge IntelligenceHeterogeneous ComputingInference Time Scaling

Operator composition and performance trade-offs remain open challenges in multi-tier edge AI deployment. Method: This paper systematically evaluates the black-box deployment efficacy of three operator classes—model partitioning, quantization, and early-exit—across mobile-edge-cloud architectures, conducting cross-device empirical studies on ResNet, ResNeXt, DUC, and FCN vision models. Contribution/Results: We first reveal that a quantization + early-exit hybrid strategy achieves optimal latency–accuracy trade-off at the edge with only marginal accuracy degradation (<2%). We propose a scenario-aware deployment decision framework grounded in resource constraints, network conditions, and input scale: (i) mobile–edge collaborative partitioning under tight resource budgets; (ii) cloud-only execution for small-scale inputs; and (iii) edge-first execution for large-scale inputs. Experimental results demonstrate that quantization alone better preserves accuracy, whereas hybrid strategies significantly improve edge energy efficiency.

Assessing latency vs accuracy trade-offs in Edge AI deployment strategiesDetermining best deployment tiers for resource-constrained Mobile/Edge scenariosEvaluating hybrid operator combinations for optimal Edge AI performance

Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical

Jul 10, 2024
AP
Adarsh Prasad Behera
🏛️ IMDEA Networks Institute | University of Helsinki | Eurecom

To address the insufficient inference capability of resource-constrained edge devices, this paper systematically compares hierarchical inference (HI) against pure on-device inference across accuracy, latency, and energy consumption. Existing HI studies overlook device-side latency and energy overhead and fail to jointly model the heterogeneous hardware, network, and model dimensions. To bridge this gap, we conduct the first multi-dimensional empirical evaluation of HI on real embedded hardware. We propose Early Exit with HI (EE-HI), a hybrid mechanism that dynamically coordinates lightweight local models and remote servers: it adaptively offloads samples and enables early exit during image classification, thereby eliminating HI’s inherent fixed overhead. Experiments show that HI reduces latency by up to 73% and device energy consumption by up to 77% compared to pure on-device inference; EE-HI further improves upon HI by reducing latency by 59.7% and device energy consumption by 60.4%.

Enhancing edge ML with hierarchical inference for complex tasksOptimizing hybrid systems for accuracy and efficiencyReducing latency and energy in on-device ML systems

Latest Papers

What's happening recently
View more

Existing edge AI inference systems are constrained by model-level mapping strategies, which hinder efficient utilization of heterogeneous computing resources to accommodate diverse operator characteristics. This work proposes the first unified operator-level scheduling framework that dynamically assigns each operator to the optimal processing unit (CPU/GPU/NPU) based on empirical performance profiling. By constructing a weighted execution graph and solving a shortest-path problem, the framework enables latency- or energy-efficiency-oriented scheduling. It transcends conventional limitations by uniformly supporting sequential execution, intra-model parallelism, and multi-model concurrency, all without relying on model-specific heuristics, thus achieving model-agnostic applicability. Experiments on an Intel Core Ultra SoC demonstrate up to 1.60× speedup with intra-model parallelism, a geometric mean acceleration of 3.42× for concurrent multi-model execution, and an average energy saving of 48.2% under energy-efficient scheduling.

edge AIheterogeneous edge inferencemodel heterogeneity

This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.

inference pipelinesmicrocontrollerneural inference

This study addresses the stringent constraints on size, power, and computational resources faced by AI inference on resource-limited platforms such as small satellites. By conducting empirical characterization of quantized AI inference on Cortex-M-class processors using representative embedded vision neural networks, the work establishes the first measurement-based performance baseline for on-board embedded systems. It introduces an explicit multi-core/multi-device cooperative scheduling mechanism and integrates analysis of ALU/SIMD utilization with memory traffic to evaluate system behavior. Moving beyond conventional paradigms that rely on opaque OS-level scheduling, this research provides comparable latency and data-movement benchmarks for typical spaceborne processors like LEON and NOEL-V, thereby demonstrating the critical role of architecture-aware design and cooperative scheduling as key dimensions in optimizing embedded AI inference for satellite applications.

embedded platformsonboard computingquantized AI inference

This work addresses the challenge of accurately predicting inference latency under dynamic voltage and frequency scaling (DVFS) on mobile edge devices, where fluctuating CPU/GPU frequencies render traditional static analysis ineffective and exhaustive empirical profiling prohibitively expensive—particularly for small language models (SLMs) with variable context lengths. To overcome this, the authors propose FLAME, a novel method that introduces the first fine-grained model of asynchronous CPU-GPU execution. FLAME enables bottom-up, frequency-aware latency prediction across both deep neural networks (DNNs) and SLMs by combining layer-level delay decomposition, quantification of parallel execution overlap and pipeline bubbles, and frequency extrapolation from sparsely sampled measurements. The approach reduces modeling time from hours or days to minutes, drastically cuts required empirical samples, maintains low prediction error, and enhances deadline-aware DVFS scheduling by jointly optimizing energy efficiency and latency guarantees, outperforming existing solutions.

asynchronous CPU-GPU couplingDynamic Voltage and Frequency Scalinglatency estimation

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
YD

Yucheng Ding

Shanghai Jiao Tong University
Device-Cloud ML
JH

Ju-Hyung Lee

Principal Researcher, Nokia
wireless communicationon-device AIdigital twins
FL

Fangming Liu

Professor, School of Computer Science & Technology, Huazhong University of Science & Technology
AI & Cloud ComputingDatacenterLLM SystemEdge Computing
ZR

Zhaochun Ren

Leiden University
Information retrievalNatural language processing