Score
Designs, builds, and evaluates machine learning inference deployments targeted to CPU-only, microcontroller-class hardware, including model compression and adaptation (e.g., quantization, pruning, architecture changes), compiler/toolchain integration, and memory/compute layout to fit flash, RAM, latency, and energy constraints. Analyzes performance, correctness, and resource tradeoffs across cross-compilation, runtime kernels, and deployment pipelines to achieve reliable, low-latency inference on embedded microcontrollers.
Deploying quantized neural networks (QNNs) on resource-constrained microcontrollers entails fundamental trade-offs among accuracy degradation, computational overhead, and memory constraints. This paper adopts a hardware–software co-design perspective to systematically analyze quantization mechanisms for embedded deployment, integrating cross-layer optimizations: model-level (e.g., asymmetric and per-channel quantization), software-level (TinyML framework adaptation and operator optimization), and hardware-level (dedicated accelerator support). We propose a QNN deployment methodology that jointly optimizes accuracy, efficiency, and framework/hardware compatibility. We empirically evaluate integration bottlenecks across mainstream TinyML frameworks—including TensorFlow Lite Micro and Apache TVM—and target hardware platforms such as ARM Cortex-M and RISC-V cores. Finally, we identify promising research directions, including ultra-low-bit quantization, structured sparsity, and compiler-aware training. Our work delivers a reproducible technical pathway and practical guidelines for robust TinyML deployment in real-world embedded systems.
This work addresses the challenge of end-to-end machine learning inference on microcontroller-class edge devices under stringent constraints on memory, energy consumption, and latency. To bridge the gap between conventional machine learning pipelines and embedded deployment realities, the authors propose a robust design framework tailored for resource-constrained environments, encompassing data acquisition, preprocessing, model compression, and streaming deployment. The framework integrates sampling buffers, feature dimensionality reduction techniques (e.g., RMS, spectral features, MFCCs), validation strategies for class imbalance, and co-optimization of models with runtime systems to form a complete embedded ML pipeline. Experimental evaluations on two representative tasks—inertial human activity recognition and keyword spotting—demonstrate that the proposed approach enables efficient, practical, and robust on-device inference, significantly narrowing the divide between general-purpose machine learning methodologies and embedded implementation requirements.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
This work addresses the limitations of existing microcontroller systems that treat AI inference merely as a conventional application library, leading to memory fragmentation, complex hardware state management, and poor control over model lifecycles. To overcome these challenges, the authors propose the first open-source runtime architecture that natively treats AI inference as a first-class workload, built upon the Zephyr RTOS. The design integrates a tensor-aware bump allocator, a four-state abstraction layer for NPU/DSP resources, a three-state model registry, and a four-tag cycle-accurate profiler. This enables zero-fragmentation tensor allocation, deterministic hardware access, secure model hot-swapping, and fine-grained performance profiling. Evaluated on the NXP MCXN947 platform, the system achieves 78,000 constant-time tensor allocations per second, an end-to-end inference overhead of only 1,038 microseconds, and 100% CI test pass rate, with all code publicly released.
To address inefficient N:M sparse DNN inference on resource-constrained microcontrollers (MCUs), this paper proposes a hardware-software co-design acceleration framework. It introduces a lightweight RISC-V instruction set extension supporting indirect loading and index decompression, develops high-efficiency sparse GEMM and convolution kernels, and integrates end-to-end sparse compilation into TVM. The key contribution is the first complete N:M sparse computing stack tailored for RISC-V MCUs, balancing hardware extensibility with software deployability. Experimental results show that the sparse kernels achieve 2.1×–3.4× speedup over dense baselines; the custom instructions further improve performance by 1.9×. End-to-end inference acceleration reaches 3.21× for ResNet-18 and 1.81× for ViT, with negligible accuracy degradation (<1.5% top-1 error).
This study addresses the stringent constraints on size, power, and computational resources faced by AI inference on resource-limited platforms such as small satellites. By conducting empirical characterization of quantized AI inference on Cortex-M-class processors using representative embedded vision neural networks, the work establishes the first measurement-based performance baseline for on-board embedded systems. It introduces an explicit multi-core/multi-device cooperative scheduling mechanism and integrates analysis of ALU/SIMD utilization with memory traffic to evaluate system behavior. Moving beyond conventional paradigms that rely on opaque OS-level scheduling, this research provides comparable latency and data-movement benchmarks for typical spaceborne processors like LEON and NOEL-V, thereby demonstrating the critical role of architecture-aware design and cooperative scheduling as key dimensions in optimizing embedded AI inference for satellite applications.
This work addresses the high energy consumption and hardware dependency of large models on resource-constrained microcontrollers by proposing an integrated compression and deployment methodology that efficiently adapts FastGRNN to 8/16-bit MCUs lacking hardware multipliers. The approach features low-rank weight decomposition, iterative hard-thresholding sparsification, Q15 post-training quantization, activation calibration, and a novel lookup-table-based acceleration mechanism for sigmoid and tanh functions tailored to multiplier-less architectures, enabling bit-accurate deterministic inference across platforms. The resulting model occupies only 566 bytes and achieves a macro F1 score of 0.918 on the HAPT dataset. It enables real-time 50 Hz inference in 9.21 ms on Arduino and 13 ms on MSP430, with the lookup-table method yielding a 30.5× speedup and reducing energy consumption by 96.7%.
This work addresses the limitations of existing Tsetlin Machine (TM) hardware, which suffers from poor programmability and low energy efficiency due to tightly coupled interfaces with external processors. Targeting edge inference scenarios, the authors propose a streamlined RISC-V-based microprocessor architecture that co-optimizes hardware and software by pruning the instruction set through instruction profiling and customizing the data path and control logic to exploit the bit-wise operations inherent to TMs. The resulting design maintains full programmability while achieving significant gains in energy efficiency: it attains accuracy comparable to or exceeding that of binary neural networks across multiple datasets—reaching 88.18% on CIFAR-2—and reduces execution time by up to 98% with an average 29.7× reduction in energy consumption.