on-device mllm deployment

Designs, builds, and evaluates systems that deploy multimodal large language models to run entirely on end-user devices, enabling local inference and on-device generation for modalities such as text, images, and speech without relying on cloud services. Implements resource-aware model optimization and runtime orchestration for constrained hardware, plus safety mechanisms (e.g., abstention), privacy-preserving behavior, and auditable interaction logging to support verifiable, offline multimodal inference.

on-devicemllmdeployment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue Settings

Feb 16, 2025
LY
Liangqi Yuan
🏛️ Purdue University | Yonsei University | IBM

To address the trade-off between resource constraints for on-device LLM deployment and high latency/cost of cloud-only execution in multimodal, multi-task, and multi-turn dialogue scenarios, this paper proposes a local-cloud collaborative inference offloading framework. Our method introduces (1) a novel Resource-Constrained Reinforcement Learning (RCRL)-driven dynamic offloading policy that jointly optimizes execution location, modality selection, and task routing; (2) M4A1—the first benchmark dataset covering multimodality, multitasking, multi-turn dialogue, and multi-LLM characteristics; and (3) a lightweight local model coupled with a large-scale cloud model, enabling multimodal input fusion and fine-grained task scheduling. Experiments on realistic multi-turn dialogues demonstrate significant reductions in end-to-end latency and cloud invocation cost, while preserving response quality—validating both effectiveness and practicality.

Addresses resource constraints in multi-modal, multi-task settings.Balances response quality, latency, and usage costs.Optimizes local-cloud inference offloading for LLMs.

Harnessing Large Language Models Locally: Empirical Results and Implications for AI PC

May 21, 2025
QS
Qingyu Song
🏛️ The Chinese University of Hong Kong | Huawei

This study addresses the joint optimization of model capability, development efficiency, and system resource constraints for local large language model (LLM) deployment on resource-constrained edge devices (e.g., AI PCs). We systematically evaluate 12 open-weight LLMs (0.5B–14B parameters) under seven post-training quantization (PTQ) schemes on CPU hardware. Our empirical analysis reveals a near-linear relationship between effective bits per weight (BPW) and inference latency, memory footprint, and power consumption—identifying ~3.5 BPW as a critical performance inflection point where low-bit large models outperform high-precision smaller ones. Quantization to low BPW incurs negligible accuracy degradation while reducing memory usage by up to 72% and shifting power consumption dominance toward compute-intensive operations. The work delivers a reproducible, empirically grounded configuration guide for edge LLM deployment, establishing an engineering paradigm that balances privacy preservation with computational efficiency.

Assessing system resource trade-offs in edge device LLM deployment.Evaluating performance limitations of on-device LLMs due to compression.Identifying optimal quantization thresholds for model efficiency and accuracy.

SmolVLM: Redefining small and efficient multimodal models

Apr 07, 2025
AM
Andrés Marafioti
🏛️ Hugging Face | Stanford University

To address the high GPU memory consumption and low inference efficiency of large vision-language models (VLMs) on mobile and edge devices, this work proposes an architecture–tokenization–data co-optimization paradigm tailored for resource-constrained scenarios. Methodologically, we design a lightweight Transformer backbone, introduce sparse image/video tokenization strategies, and construct a high-quality, compact multimodal dataset trained via curriculum learning. Our key contributions are: (1) SmolVLM-256M achieves <1 GB GPU memory usage during inference while outperforming Idefics-80B in accuracy; (2) the 2.2B-parameter variant attains state-of-the-art performance on both image and video understanding tasks with significantly lower memory footprint; and (3) this is the first demonstration of a small-parameter VLM systematically surpassing ultra-large models across multimodal understanding benchmarks—establishing a new paradigm for efficient VLM deployment on edge devices.

Develop compact multimodal models for resource-efficient inferenceEnhance performance with minimal GPU memory usageOptimize architectural configurations for low computational overhead

Institutional Platform for Secure Self-Service Large Language Model Exploration

Feb 01, 2024
VK
V. K. Cody Bumgardner
🏛️ University of Kentucky

This work addresses the challenge of securely, efficiently, and accessibly deploying large-scale, customized large language models (LLMs) in multi-tenant academic research environments. We propose a tenant-aware, proxy-based computational network architecture that integrates zero-trust principles, secure sandboxing, role-based fine-grained access control (RBAC), and end-to-end HTTPS/TLS encryption to enable logically unified scheduling of physical resource pools while enforcing strict inter-tenant process and data isolation. Our approach innovatively combines parallel multi-LoRA inference, agent-driven resource orchestration, and an encrypted inference pipeline. Deployed at the University of Kentucky’s AI Center, the platform supports secure, cross-disciplinary LLM usage by research teams. Evaluation shows a 37% reduction in inference latency and near-zero inter-tenant resource leakage risk, demonstrating robust security, scalability, and operational efficiency in production academic settings.

Enhances accessibility to large language modelsFacilitates secure, isolated AI resource utilizationSecures multi-LoRA inference for diverse users

SPA: Towards A Computational Friendly Cloud-Base and On-Devices Collaboration Seq2seq Personalized Generation

Mar 11, 2024
YL
Yanming Liu
🏛️ Zhejiang University | Southeast University | Massachusetts Institute of Technology | Beijing Institute of Technology | Tianjin University | Tongji University

To address the high memory footprint and slow inference of large language models (LLMs) on resource-constrained edge devices, this paper proposes the lightweight Side-Plugin Adaptation (SPA) architecture for efficient cloud-edge collaborative seq2seq generation. Our method introduces a novel cloud-edge parameter decoupling paradigm: general knowledge is retained in the cloud-based LLM, while personalized parameters are deployed on-device as ultra-lightweight side plugins—supporting hierarchical deployment and incremental fine-tuning. Under stringent memory and computational constraints, SPA ensures fast, stable on-device inference. Experimental results demonstrate that SPA reduces inference latency by 42% and memory consumption by 67% compared to state-of-the-art on-device LLM approaches, significantly enhancing both personalization efficiency and practical deployability.

Enabling cloud-device collaboration for personalized generationImproving computational speed on constrained devicesReducing memory storage for LLMs on low-resource devices

Latest Papers

What's happening recently
View more

This work addresses the challenge of scheduling heterogeneous requests—comprising video, image, and text—in multimodal large language model inference, where significant disparities in resource demands cause head-of-line blocking and severe latency spikes under conventional text-oriented schedulers. To overcome this, the authors propose RPS-Serve, a modality-aware scheduler that abstracts incoming requests as “rocks” (video), “pebbles” (image), and “sand” (text). By integrating dynamic priority assignment with an aging mechanism, RPS-Serve enables lightweight requests to swiftly bypass heavy workloads without compromising fairness. Evaluated on mainstream multimodal large models, the approach reduces average time-to-first-token latency by 54% and cuts latency for delay-sensitive requests by 78.5%, substantially improving both system throughput and resource efficiency.

Head-of-Line BlockingInference SchedulingLatency

Existing multimodal inference systems face challenges including rigid workflow orchestration, inefficient intermediate data transfer, and difficulty in sharing KV caches and model weights across heterogeneous components. This work proposes the first system-level unified abstraction for multimodal inference, featuring a three-tier architecture that co-optimizes control flow, data flow, and compute flow. The control flow layer employs a Python DSL to support both static and dynamic workflow orchestration; the data flow layer implements a zero-copy, distributed paged KV cache spanning GPUs, CPUs, and SSDs; and the compute flow layer enables multimodal prefix matching and KV reuse, unifying the forward passes of LLMs and diffusion models through a common SGLang interface. This design decouples orchestration logic from data transmission mechanisms, achieving efficient inference and resource reuse across diverse scenarios such as LongCat-Next dialogue and HunyuanImage-3 generation.

distributed KV cacheheterogeneous computingintermediate data transmission

This work addresses the high energy consumption of large language model (LLM) inference and the inefficiency of conventional parameter tuning methods, which often require days of computation and struggle to adapt to diverse hardware and system constraints. The authors propose a human-in-the-loop optimization framework that integrates conversational LLMs with human feedback, leveraging an enhanced prompt template to enable rapid, adaptive search over inference runtime parameters. Their approach significantly improves tuning efficiency, converging to configurations below a target energy threshold in just 3.4 prompts on average—outperforming baseline methods such as Sobol sampling, which requires 5.2 prompts—and consistently achieves lower energy consumption per token, demonstrating clear advantages in both convergence speed and energy efficiency.

Energy EfficiencyLarge Language ModelsModel Inference

This paper identifies “modality inflation”—a phenomenon in multimodal large language model (MLLM) inference where visual inputs induce substantial computational and energy overhead due to redundant visual encoding and excessively long visual token sequences. Method: Leveraging fine-grained power profiling on NVIDIA A100 GPUs, we conduct the first energy-efficiency attribution analysis across prefill, decoding, and visual encoding stages, revealing 17%–94% higher energy consumption versus text-only baselines, with bottlenecks varying heterogeneously across models (either in visual encoding or long-sequence processing). We propose a stage-adaptive Dynamic Voltage and Frequency Scaling (DVFS) strategy that achieves significant energy savings under latency constraints, incurring <8% throughput degradation. Contribution/Results: We formally define “modality inflation,” establish the first MLLM-specific energy-attribution methodology, and empirically validate the efficacy of architecture-aware, stage-specific energy optimization.

Analyzes energy inefficiency in multimodal large language modelsProposes dynamic voltage scaling for energy-efficient MLLM inferenceQuantifies energy overhead from vision encoding and token expansion

This work addresses the severe memory capacity and bandwidth bottlenecks faced by generative AI during inference on resource-constrained devices, particularly in long-context and multimodal scenarios. The authors propose a hierarchical roofline performance model to systematically evaluate, for the first time, the bandwidth and latency requirements of high-bandwidth storage (HBS) in large-model long-context inference, establishing clear HBS performance thresholds necessary to achieve interactive throughput. For smaller models, they design an efficient memory utilization scheme leveraging bonded global buffer chips. Experimental results demonstrate that the proposed approaches substantially alleviate memory pressure and improve energy efficiency, offering a critical technical pathway for deploying generative AI at the edge.

generative AI inferenceKey-Value cachinglong context lengths

Hot Scholars

PZ

Pengfei Zhao

ATB Potsdam
LLMCompressionXAIMechanistic Interpretability
LB

Lidong Bing

MiroMind, Alibaba DAMO, Tencent, CMU, CUHK
Natural Language ProcessingLarge Language ModelsLarge Multimodal Models
JD

Jiankang Deng

Imperial College London
Computer VisionMachine Learning
WC

Weidong Cai

Clinical Associate Professor, Stanford University School of Medicine
functional neuroimagingmachine learningcognitivedevelopmental
YZ

Yueyi Zhang

Miromind, Previously University of Science and Technology of China
Structured lightDepth SensingEvent CameraMedical Imaging