Hardware optimization on Android for inference of AI models

📅 2025-11-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of efficiently orchestrating heterogeneous hardware (e.g., GPU, NPU) for real-time AI inference on Android devices, this paper proposes a mobile-oriented, hardware-aware inference optimization framework. Methodologically, it integrates model quantization (INT8/FP16), operator-level hardware mapping, and coordinated accelerator scheduling to derive optimal execution configurations for YOLO (object detection) and ResNet (image classification) models. Its key contributions are: (1) the first system-level joint optimization of quantization accuracy, hardware resource utilization, and inference latency on Android; and (2) a lightweight configuration search mechanism enabling rapid cross-platform adaptation across Qualcomm, MediaTek, and Huawei NPUs. Experiments demonstrate that, with <1.2% mAP/Top-1 accuracy degradation, the framework achieves an average 58.3% reduction in end-to-end inference latency and a 2.1× improvement in energy efficiency—significantly outperforming TensorFlow Lite and ONNX Runtime mobile deployments.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Learning & Optimization for CVSearch and Optimization: Algorithm Configuration

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
The pervasive integration of Artificial Intelligence models into contemporary mobile computing is notable across numerous use cases, from virtual assistants to advanced image processing. Optimizing the mobile user experience involves minimal latency and high responsiveness from deployed AI models with challenges from execution strategies that fully leverage real time constraints to the exploitation of heterogeneous hardware architecture. In this paper, we research and propose the optimal execution configurations for AI models on an Android system, focusing on two critical tasks: object detection (YOLO family) and image classification (ResNet). These configurations evaluate various model quantization schemes and the utilization of on device accelerators, specifically the GPU and NPU. Our core objective is to empirically determine the combination that achieves the best trade-off between minimal accuracy degradation and maximal inference speed-up.
Problem

Research questions and friction points this paper is trying to address.

Optimizing AI model inference latency on Android hardware
Evaluating quantization and accelerator usage for mobile AI
Balancing accuracy preservation with inference speed acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimizes AI model execution on Android
Evaluates quantization schemes and hardware accelerators
Balances accuracy and speed for mobile inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Iulius Gherasim
Department of Computer Architecture and Automation, Complutense University of Madrid, Madrid, Spain
Carlos García Sánchez
Carlos García Sánchez
Department of Computer Architecture and Automation, Complutense University of Madrid, Madrid, Spain