Score
Designs and builds end-to-end local model deployment systems and pipelines that implement efficient on-device and embedded inference, including frozen-model integration, hybrid or joint-model setups, lightweight and metamodel architectures, and finite/optimized model implementations. Implements and analyzes optimized inference engines and deployment workflows — covering frozen-model interfacing, incremental and continuous test-time updates, and large-scale evaluation of latency, throughput, memory, and energy trade-offs to balance accuracy and efficiency.
Deploying complex AI models efficiently on resource-constrained edge devices (e.g., smartphones, IoT endpoints) remains challenging due to stringent latency, memory, and energy constraints. Method: This paper proposes a three-tier co-optimization paradigm—spanning data, model, and system layers—unifying cross-layer coordination mechanisms for the first time and bridging a critical methodological gap in end-to-cloud full-stack optimization. It integrates lightweight edge data preprocessing, model compression techniques (pruning, quantization, knowledge distillation), efficient inference runtimes (TensorFlow Lite, ONNX Runtime), and heterogeneous hardware acceleration (NPU/GPU/FPGA). Contribution/Results: We construct a reusable edge AI optimization roadmap, achieving an average 3.2× inference speedup and 78% memory reduction while preserving ≥95% of the original model accuracy.
This study addresses the challenge of simultaneously achieving low latency, high throughput, and cost efficiency in enterprise-scale composite AI systems under concurrent heterogeneous model invocations. It presents the first systematic analysis of system-specific issues, including multi-model fan-out overhead, cascading cold-start propagation, and dynamic heterogeneity in scaling behavior. To tackle these challenges, the authors propose a modular, platform-agnostic inference architecture that integrates serverless computing, dynamic autoscaling, MLOps pipelines, and cooperative multi-model scheduling. Experimental results demonstrate that the proposed approach reduces P95 tail latency by over 50%, increases throughput by up to 3.9×, and achieves 30%–40% cost savings compared to baseline systems.
To address low testing and debugging efficiency and immature toolchains when deploying rapidly evolving large language models (LLMs) on emerging platforms (e.g., browsers, mobile devices), this paper proposes TapML—a top-down, test-driven framework. Methodologically, TapML introduces (1) the first operator-level test pruning technique that automatically generates high-coverage, realistic test inputs; (2) a progressive cross-platform migration strategy that significantly narrows the scope for compound error localization; and (3) native backend support for Metal and WebGPU, with deep integration into MLC-LLM. Evaluated over two years, TapML has enabled efficient deployment of 105 emerging models—spanning 27 distinct architectures—across five platform categories, reducing average deployment time by 42%. It has since become the default development paradigm for MLC-LLM.
Active Inference (AIF) suffers from high computational and memory overhead, hindering its deployment on resource-constrained real-time or embedded systems. To address this, we propose a hardware-efficient AIF computing architecture: leveraging the pymdp framework, we construct a sparse, unified computational graph that explicitly optimizes computation flow and memory access patterns while preserving model flexibility. This work presents the first customized hardware adaptation of AIF for edge devices. Experimental evaluation demonstrates substantial efficiency gains—over 2× latency reduction and up to 35% lower peak memory footprint—without compromising inference fidelity. Our core innovation lies in mapping AIF’s Bayesian inference process onto a sparse, statically schedulable computational graph, thereby jointly optimizing algorithmic accuracy and hardware execution efficiency. The proposed architecture provides a scalable, system-level solution for deploying lightweight active agents in practical edge scenarios.
This work addresses the cold-start latency challenges faced by vLLM in large-scale inference serving, a problem whose root causes have not been systematically investigated. The study presents the first fine-grained decomposition of vLLM’s cold-start process, identifying six critical stages and revealing their CPU-bound nature. Building upon this insight, the authors develop an interpretable latency attribution and prediction model that integrates both model-level and system-level parameters. By leveraging performance profiling, parametric modeling, and optimizations such as torch.compile, the proposed approach achieves high-accuracy prediction of startup latency across diverse hardware configurations. The project publicly releases the complete suite of analysis tools and datasets, offering practical foundations for resource provisioning and performance optimization in inference serving systems.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
This work addresses the challenges of automating the deployment of open-source AI models as callable APIs, a process often hindered by complexity and prolonged timelines. To overcome the inefficiencies of existing approaches, the authors propose MADE, a novel system featuring a belief-driven dual-agent collaborative architecture that iteratively constructs, validates, and retrospectively refines deployment artifacts through execution feedback, enabling end-to-end autonomous deployment. Key contributions include a large language model–based mechanism for dual-agent coordination and belief updating, an automated pipeline for code generation and test validation, a method for parsing heterogeneous model resources, and M2ABench—the first benchmark for model-to-API conversion. Evaluated on M2ABench, which comprises 122 real-world models, MADE achieves a deployment success rate of 68.85%, substantially outperforming SWE-agent and OpenHands.
This study addresses the lack of systematic analysis of neural processing unit (NPU) efficiency at both operator and pipeline levels in on-device large language model (LLM) inference. The authors propose an OPMASK-driven controlled pipeline decomposition approach, integrating stage-level performance profiling with fine-grained energy and latency measurements on heterogeneous SoCs, thereby revealing for the first time the critical bottlenecks in CPU-NPU collaborative inference. Experimental results demonstrate that the CPU can be up to 1.6× faster during the Prefill stage, while NPU acceleration in the Decode stage yields only modest speedups of 1.05–1.2×. Moreover, due to scheduling overhead and cross-backend fallbacks, naively offloading computation to the NPU can increase energy consumption by up to 51%. This work provides novel insights and methodological foundations for efficient LLM deployment on mobile devices.
This work addresses the high inference costs, low service efficiency, and insufficient stability of large language models by proposing the first token-centric four-layer inference optimization framework. The architecture integrates multi-model fusion, model compression and quantization, compute-model co-optimization, and joint scheduling across computation, networking, and modeling. By systematically combining these key techniques, the framework substantially reduces the cost per generated token while significantly enhancing service efficiency and supply stability. It provides a holistic, efficient, stable, and cost-effective solution that enables large models to transition from being merely callable to truly operable at scale, thereby supporting their widespread deployment in real-world applications.