on-device multimodal inference

Design and build inference systems that run and fuse multiple data modalities locally on mobile or edge devices, implementing sensor I/O, modality fusion, and runtime integration. Optimize and analyze models and pipelines for constrained hardware by applying compression and acceleration techniques (quantization, pruning, operator selection), selecting or integrating lightweight runtimes, and balancing latency, memory, power, accuracy, and on‑device privacy.

on-devicemultimodalinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models

Mar 08, 2025
XW
Xubin Wang
🏛️ Hong Kong Baptist University | BNU-HKBU United International College | Beijing Normal University

This paper addresses core challenges in deploying AI models on edge and end devices—namely, poor real-time performance, severe resource constraints, and data privacy risks. To this end, it proposes the first unified definition and multidimensional analytical framework for on-device AI, integrating dual perspectives from edge computing and large language model evolution, and establishing a systematic “end-to-end challenges–technical solutions” mapping. Methodologically, it comprehensively synthesizes key techniques including model compression (pruning, quantization, distillation), lightweight architecture design, hardware-software co-optimization, edge-side preprocessing, and federated learning. Contributions include: (1) categorizing 12 representative application scenarios; (2) identifying 7 fundamental technical bottlenecks; and (3) distilling 4 practical, deployable evolutionary pathways. The resulting structured survey serves as an authoritative benchmark and implementation guide for both industrial deployment and academic research.

Discusses optimization strategies for on-device AI implementationExamines impact of emerging technologies on edge AI evolutionExplores challenges of deploying AI models on edge devices

Must-Read Papers

Most classic and influential ideas
View more

Hardware optimization on Android for inference of AI models

Nov 17, 2025
IG
Iulius Gherasim
🏛️ Complutense University of Madrid

To address the challenge of efficiently orchestrating heterogeneous hardware (e.g., GPU, NPU) for real-time AI inference on Android devices, this paper proposes a mobile-oriented, hardware-aware inference optimization framework. Methodologically, it integrates model quantization (INT8/FP16), operator-level hardware mapping, and coordinated accelerator scheduling to derive optimal execution configurations for YOLO (object detection) and ResNet (image classification) models. Its key contributions are: (1) the first system-level joint optimization of quantization accuracy, hardware resource utilization, and inference latency on Android; and (2) a lightweight configuration search mechanism enabling rapid cross-platform adaptation across Qualcomm, MediaTek, and Huawei NPUs. Experiments demonstrate that, with <1.2% mAP/Top-1 accuracy degradation, the framework achieves an average 58.3% reduction in end-to-end inference latency and a 2.1× improvement in energy efficiency—significantly outperforming TensorFlow Lite and ONNX Runtime mobile deployments.

Balancing accuracy preservation with inference speed accelerationEvaluating quantization and accelerator usage for mobile AIOptimizing AI model inference latency on Android hardware

Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical

Jul 10, 2024
AP
Adarsh Prasad Behera
🏛️ IMDEA Networks Institute | University of Helsinki | Eurecom

To address the insufficient inference capability of resource-constrained edge devices, this paper systematically compares hierarchical inference (HI) against pure on-device inference across accuracy, latency, and energy consumption. Existing HI studies overlook device-side latency and energy overhead and fail to jointly model the heterogeneous hardware, network, and model dimensions. To bridge this gap, we conduct the first multi-dimensional empirical evaluation of HI on real embedded hardware. We propose Early Exit with HI (EE-HI), a hybrid mechanism that dynamically coordinates lightweight local models and remote servers: it adaptively offloads samples and enables early exit during image classification, thereby eliminating HI’s inherent fixed overhead. Experiments show that HI reduces latency by up to 73% and device energy consumption by up to 77% compared to pure on-device inference; EE-HI further improves upon HI by reducing latency by 59.7% and device energy consumption by 60.4%.

Enhancing edge ML with hierarchical inference for complex tasksOptimizing hybrid systems for accuracy and efficiencyReducing latency and energy in on-device ML systems

Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing

Jun 03, 2025
JS
Jinhao Sheng
🏛️ Beijing Normal University | BNU-HKBU United International College

This paper addresses the deployment of multi-model inference pipelines on resource-constrained edge devices. We propose an end-to-end adaptive configuration framework that, for the first time, explicitly incorporates device resource constraints into joint optimization decisions. Our method integrates residual feature extraction, LSTM-based workload forecasting, and policy-gradient reinforcement learning to jointly optimize QoS guarantees (e.g., latency and throughput), operational cost, and real-time adaptability. Evaluated on a real Kubernetes-based edge cluster, the framework achieves a 27% reduction in average inference latency, a 31% increase in throughput, a 22% decrease in deployment cost, and over 42% faster configuration decision-making for complex pipelines—outperforming state-of-the-art baselines. The core contributions are: (i) a resource-aware joint optimization model that unifies hardware constraints with pipeline scheduling and scaling decisions; and (ii) a lightweight, learning-driven configuration mechanism enabling efficient, online adaptation under dynamic edge conditions.

Address device resource constraints in pipeline configurationOptimize QoS and costs for edge multi-model inference pipelinesReduce decision-making time for complex edge pipelines

This work addresses the redundant computation bottleneck in feature extraction during on-device model inference by formally modeling the feature extraction pipeline as a directed acyclic graph (DAG) and introducing graph optimization techniques to enable cross-feature fusion and cross-inference caching. The proposed approach effectively eliminates redundant computations both across features and between consecutive inference requests, significantly accelerating the feature preparation phase prior to on-device inference without compromising model accuracy. Deployment and validation across five industrial-scale mobile services—including search, video, and e-commerce—demonstrate substantial end-to-end latency reductions: 1.33×–3.93× during daytime and 1.43×–4.53× at night.

feature extractionlatency optimizationon-device inference

Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power, thermal envelope, and memory. We benchmark Qwen 2.5 1.5B (4-bit quantised) across four platforms: a Raspberry Pi 5 with Hailo-10H NPU, a Samsung Galaxy S24 Ultra, an iPhone 16 Pro, and a laptop NVIDIA RTX 4050 GPU. Using a fixed 258-token prompt over 20 warm-condition iterations per device, we measure throughput, latency, power, and thermal behaviour. For mobile platforms, thermal management supersedes peak compute as the primary constraint: the iPhone 16 Pro loses nearly half its throughput within two iterations, and the S24 Ultra suffers a hard OS-enforced GPU frequency floor that terminates inference entirely. On dedicated hardware, distinct constraints dominate: the RTX 4050 is bounded by its battery power ceiling, while the Hailo-10H is limited by on-module memory bandwidth. The RTX 4050 sustains 131.7 tok/s at 34.1 W; the Hailo-10H sustains 6.9 tok/s at under 2 W with near-zero variance, matching the RTX 4050 in energy proportionality at 19x lower throughput. Results should be interpreted as platform-level deployment characterisations for a single model and prompt type, reflecting hardware and software combined, rather than general claims about hardware capability alone.

edge computinghardware limitationsLLM inference

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic analysis of neural processing unit (NPU) efficiency at both operator and pipeline levels in on-device large language model (LLM) inference. The authors propose an OPMASK-driven controlled pipeline decomposition approach, integrating stage-level performance profiling with fine-grained energy and latency measurements on heterogeneous SoCs, thereby revealing for the first time the critical bottlenecks in CPU-NPU collaborative inference. Experimental results demonstrate that the CPU can be up to 1.6× faster during the Prefill stage, while NPU acceleration in the Decode stage yields only modest speedups of 1.05–1.2×. Moreover, due to scheduling overhead and cross-backend fallbacks, naively offloading computation to the NPU can increase energy consumption by up to 51%. This work provides novel insights and methodological foundations for efficient LLM deployment on mobile devices.

energy consumptionheterogeneous executionmobile LLM inference

This work addresses the challenge of on-device model adaptation under resource constraints, where end-to-end backpropagation through modern deep networks is infeasible on edge hardware. The authors propose a heterogeneous adaptation pipeline that deploys an INT8-quantized backbone on a Hailo-8L inference accelerator for frozen feature extraction, while fine-tuning only a lightweight classification head in FP32 on the CPU. A post-training quantization recovery mechanism is introduced to preserve feature quality. This approach uniquely repurposes commercial inference accelerators for on-device training and overcomes training efficiency bottlenecks through co-design of the backbone and head. Experiments demonstrate up to a 15.4× speedup in training over a Raspberry Pi 5 CPU baseline, substantially reduced per-sample energy consumption, and competitive accuracy across multiple architectures and datasets.

backpropagationedge AImodel personalization

This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.

inference pipelinesmicrocontrollerneural inference

This work addresses the severe memory capacity and bandwidth bottlenecks faced by generative AI during inference on resource-constrained devices, particularly in long-context and multimodal scenarios. The authors propose a hierarchical roofline performance model to systematically evaluate, for the first time, the bandwidth and latency requirements of high-bandwidth storage (HBS) in large-model long-context inference, establishing clear HBS performance thresholds necessary to achieve interactive throughput. For smaller models, they design an efficient memory utilization scheme leveraging bonded global buffer chips. Experimental results demonstrate that the proposed approaches substantially alleviate memory pressure and improve energy efficiency, offering a critical technical pathway for deploying generative AI at the edge.

generative AI inferenceKey-Value cachinglong context lengths

Hot Scholars

ZY

Zhiyuan Yu

Assistant Professor in CSE, Texas A&M University
WL

Weiqing Luo

Arizona State University
multimodal large language modelinformation retrievalrecommendation system
XW

Xiao Wang

Professor, Arizona State University
synthetic biologysystems biology
ZH

Ziyi Huang

Assistant Professor @ Arizona State University
Trustworthy AI for Health