Score
Designs, builds, and evaluates systems that deploy multimodal large language models to run entirely on end-user devices, enabling local inference and on-device generation for modalities such as text, images, and speech without relying on cloud services. Implements resource-aware model optimization and runtime orchestration for constrained hardware, plus safety mechanisms (e.g., abstention), privacy-preserving behavior, and auditable interaction logging to support verifiable, offline multimodal inference.
Deploying large language models (LLMs) on edge devices faces fundamental challenges including constrained computational resources, limited memory capacity, and hardware heterogeneity. To address these, this paper systematically surveys the full lifecycle of edge LLMs and introduces a comprehensive, stack-wide technical taxonomy—spanning lightweight model design (e.g., pruning, quantization, distillation), runtime optimization (e.g., memory-aware inference, device-adaptive scheduling), and on-device deployment (e.g., cloud-edge collaborative frameworks). It proposes a novel cross-platform co-deployment paradigm that unifies pre-deployment model compression with dynamic execution optimization. Based on a rigorous synthesis of over 120 state-of-the-art studies, the work identifies five persistent technical bottlenecks and six key future research directions. The resulting methodology provides both a reusable conceptual framework and practical guidelines for deploying AI at the edge.
To address the trade-off between resource constraints for on-device LLM deployment and high latency/cost of cloud-only execution in multimodal, multi-task, and multi-turn dialogue scenarios, this paper proposes a local-cloud collaborative inference offloading framework. Our method introduces (1) a novel Resource-Constrained Reinforcement Learning (RCRL)-driven dynamic offloading policy that jointly optimizes execution location, modality selection, and task routing; (2) M4A1—the first benchmark dataset covering multimodality, multitasking, multi-turn dialogue, and multi-LLM characteristics; and (3) a lightweight local model coupled with a large-scale cloud model, enabling multimodal input fusion and fine-grained task scheduling. Experiments on realistic multi-turn dialogues demonstrate significant reductions in end-to-end latency and cloud invocation cost, while preserving response quality—validating both effectiveness and practicality.
This study addresses the joint optimization of model capability, development efficiency, and system resource constraints for local large language model (LLM) deployment on resource-constrained edge devices (e.g., AI PCs). We systematically evaluate 12 open-weight LLMs (0.5B–14B parameters) under seven post-training quantization (PTQ) schemes on CPU hardware. Our empirical analysis reveals a near-linear relationship between effective bits per weight (BPW) and inference latency, memory footprint, and power consumption—identifying ~3.5 BPW as a critical performance inflection point where low-bit large models outperform high-precision smaller ones. Quantization to low BPW incurs negligible accuracy degradation while reducing memory usage by up to 72% and shifting power consumption dominance toward compute-intensive operations. The work delivers a reproducible, empirically grounded configuration guide for edge LLM deployment, establishing an engineering paradigm that balances privacy preservation with computational efficiency.
To address the high GPU memory consumption and low inference efficiency of large vision-language models (VLMs) on mobile and edge devices, this work proposes an architecture–tokenization–data co-optimization paradigm tailored for resource-constrained scenarios. Methodologically, we design a lightweight Transformer backbone, introduce sparse image/video tokenization strategies, and construct a high-quality, compact multimodal dataset trained via curriculum learning. Our key contributions are: (1) SmolVLM-256M achieves <1 GB GPU memory usage during inference while outperforming Idefics-80B in accuracy; (2) the 2.2B-parameter variant attains state-of-the-art performance on both image and video understanding tasks with significantly lower memory footprint; and (3) this is the first demonstration of a small-parameter VLM systematically surpassing ultra-large models across multimodal understanding benchmarks—establishing a new paradigm for efficient VLM deployment on edge devices.
This work addresses the challenge of securely, efficiently, and accessibly deploying large-scale, customized large language models (LLMs) in multi-tenant academic research environments. We propose a tenant-aware, proxy-based computational network architecture that integrates zero-trust principles, secure sandboxing, role-based fine-grained access control (RBAC), and end-to-end HTTPS/TLS encryption to enable logically unified scheduling of physical resource pools while enforcing strict inter-tenant process and data isolation. Our approach innovatively combines parallel multi-LoRA inference, agent-driven resource orchestration, and an encrypted inference pipeline. Deployed at the University of Kentucky’s AI Center, the platform supports secure, cross-disciplinary LLM usage by research teams. Evaluation shows a 37% reduction in inference latency and near-zero inter-tenant resource leakage risk, demonstrating robust security, scalability, and operational efficiency in production academic settings.
To address the high memory footprint and slow inference of large language models (LLMs) on resource-constrained edge devices, this paper proposes the lightweight Side-Plugin Adaptation (SPA) architecture for efficient cloud-edge collaborative seq2seq generation. Our method introduces a novel cloud-edge parameter decoupling paradigm: general knowledge is retained in the cloud-based LLM, while personalized parameters are deployed on-device as ultra-lightweight side plugins—supporting hierarchical deployment and incremental fine-tuning. Under stringent memory and computational constraints, SPA ensures fast, stable on-device inference. Experimental results demonstrate that SPA reduces inference latency by 42% and memory consumption by 67% compared to state-of-the-art on-device LLM approaches, significantly enhancing both personalization efficiency and practical deployability.
This work addresses the challenge of scheduling heterogeneous requests—comprising video, image, and text—in multimodal large language model inference, where significant disparities in resource demands cause head-of-line blocking and severe latency spikes under conventional text-oriented schedulers. To overcome this, the authors propose RPS-Serve, a modality-aware scheduler that abstracts incoming requests as “rocks” (video), “pebbles” (image), and “sand” (text). By integrating dynamic priority assignment with an aging mechanism, RPS-Serve enables lightweight requests to swiftly bypass heavy workloads without compromising fairness. Evaluated on mainstream multimodal large models, the approach reduces average time-to-first-token latency by 54% and cuts latency for delay-sensitive requests by 78.5%, substantially improving both system throughput and resource efficiency.
Existing multimodal inference systems face challenges including rigid workflow orchestration, inefficient intermediate data transfer, and difficulty in sharing KV caches and model weights across heterogeneous components. This work proposes the first system-level unified abstraction for multimodal inference, featuring a three-tier architecture that co-optimizes control flow, data flow, and compute flow. The control flow layer employs a Python DSL to support both static and dynamic workflow orchestration; the data flow layer implements a zero-copy, distributed paged KV cache spanning GPUs, CPUs, and SSDs; and the compute flow layer enables multimodal prefix matching and KV reuse, unifying the forward passes of LLMs and diffusion models through a common SGLang interface. This design decouples orchestration logic from data transmission mechanisms, achieving efficient inference and resource reuse across diverse scenarios such as LongCat-Next dialogue and HunyuanImage-3 generation.
This work addresses the high energy consumption of large language model (LLM) inference and the inefficiency of conventional parameter tuning methods, which often require days of computation and struggle to adapt to diverse hardware and system constraints. The authors propose a human-in-the-loop optimization framework that integrates conversational LLMs with human feedback, leveraging an enhanced prompt template to enable rapid, adaptive search over inference runtime parameters. Their approach significantly improves tuning efficiency, converging to configurations below a target energy threshold in just 3.4 prompts on average—outperforming baseline methods such as Sobol sampling, which requires 5.2 prompts—and consistently achieves lower energy consumption per token, demonstrating clear advantages in both convergence speed and energy efficiency.
This paper identifies “modality inflation”—a phenomenon in multimodal large language model (MLLM) inference where visual inputs induce substantial computational and energy overhead due to redundant visual encoding and excessively long visual token sequences. Method: Leveraging fine-grained power profiling on NVIDIA A100 GPUs, we conduct the first energy-efficiency attribution analysis across prefill, decoding, and visual encoding stages, revealing 17%–94% higher energy consumption versus text-only baselines, with bottlenecks varying heterogeneously across models (either in visual encoding or long-sequence processing). We propose a stage-adaptive Dynamic Voltage and Frequency Scaling (DVFS) strategy that achieves significant energy savings under latency constraints, incurring <8% throughput degradation. Contribution/Results: We formally define “modality inflation,” establish the first MLLM-specific energy-attribution methodology, and empirically validate the efficacy of architecture-aware, stage-specific energy optimization.
This work addresses the severe memory capacity and bandwidth bottlenecks faced by generative AI during inference on resource-constrained devices, particularly in long-context and multimodal scenarios. The authors propose a hierarchical roofline performance model to systematically evaluate, for the first time, the bandwidth and latency requirements of high-bandwidth storage (HBS) in large-model long-context inference, establishing clear HBS performance thresholds necessary to achieve interactive throughput. For smaller models, they design an efficient memory utilization scheme leveraging bonded global buffer chips. Experimental results demonstrate that the proposed approaches substantially alleviate memory pressure and improve energy efficiency, offering a critical technical pathway for deploying generative AI at the edge.