embedded inference integration

Design and implement software artifacts and build processes that embed trained machine learning models into constrained runtime environments, including generating deterministic C modules and compiling inference into application binaries to remove external ML runtime dependencies. Optimize, validate, and integrate the inference code for platform constraints and low-latency execution (for example microsecond-scale inference) while ensuring deterministic behavior.

embeddedinferenceintegration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$275K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning

Sep 13, 2025
RM
Roberto Morabito
🏛️ EURECOM | University of Helsinki

This paper addresses the prevalent “silent failures” and unpredictable behavior of large language models (LLMs) when auto-generating code for embedded machine learning (ML) workflows. We propose a closed-loop evaluation framework covering data preprocessing, model conversion, and on-device inference code generation. Through multi-model empirical analysis, we introduce the first failure taxonomy for LLM-generated code in embedded ML, identifying systemic fragility arising from prompt format bias, implicit structural assumptions encoded in LLMs, and blind spots in compilation- and runtime-level validation. Key failure patterns include format-misleading parsing errors and “compilable-yet-functionally-broken” runtime errors—both largely undetectable by conventional verification methods. Our findings provide both theoretical foundations and practical guidelines for enhancing the reliability, traceability, and robustness of LLM-driven embedded ML systems.

Analyzing how prompt format and model assumptions cause silent failuresIdentifying error-prone behaviors that evade standard validation methodsInvestigating failure modes in LLM-powered embedded ML pipelines

Deploying large code language models on resource-constrained local devices remains challenging due to hardware limitations, which compromises privacy preservation, inference latency, and offline availability. This work proposes Ditto, a method that co-optimizes model compression and inference program generation to compile large code models into lightweight executables suitable for statically typed languages such as C. Ditto innovatively integrates bounded-error product quantization with quantization-aware inference program synthesis and extends the LLVM compiler to automatically replace general matrix-vector (GEMV) operations with efficient BLAS calls. Evaluated across three prominent code language models, Ditto achieves up to 10.5× speedup and 6.4× reduction in memory footprint, with an average drop of only 0.27% in pass@1 accuracy.

Code LLMslocal deploymentmodel efficiency

This work addresses the high code complexity and maintenance overhead commonly faced by general-purpose large language model (LLM) inference frameworks due to their need to support diverse models, hardware platforms, and optimization strategies. The authors propose a novel “LLM-as-compiler” paradigm that uniquely integrates explicit knowledge-driven reasoning with multi-agent collaboration, establishing a closed-loop pipeline of generation, verification, and knowledge distillation. In this framework, users specify only high-level inference constraints, and the system—leveraging a Contractual Knowledge Base (CKB) and coordinated multi-agent interaction—automatically synthesizes lightweight, redundancy-free, customized inference engines. Experimental results demonstrate that the approach can successfully generate high-performance, executable inference engines without any reference implementation, thereby validating the feasibility and effectiveness of knowledge-guided automated construction of efficient LLM inference systems.

code complexityLLM inferencemaintenance cost

Large language model inference lacks output determinism due to floating-point non-associativity, dynamic batching, and varying GPU reduction orders. This work proposes a scheduling-based speculative validation mechanism that introduces speculative execution into deterministic inference for the first time. By employing lightweight validate-and-rollback cycles combined with fixed-shape reduction scheduling, the approach incurs overhead only for requests requiring determinism, while remaining compatible with dynamic batching and requiring minimal modification to existing GPU kernels. The method decouples determinism guarantees from low-level implementation details, achieving high throughput and significantly outperforming baseline strategies such as disabling dynamic batching or rewriting kernel functions.

dynamic batchingfloating-point non-associativityGPU kernels

Large Language Model powered Symbolic Execution

Apr 02, 2025
YL
Yihe Li
🏛️ National University of Singapore

Existing large language models (LLMs) face significant challenges in direct symbolic execution—including low precision, high computational overhead, and strong dependence on large-scale models and high-end hardware—hindering practical deployment. To address these limitations, we propose a novel LLM-driven lightweight symbolic execution paradigm. Our approach employs path-guided task decomposition to decouple complex program analysis into fine-grained, resource-efficient subtasks. We introduce the first path-constraint generalization method based on universal code representations—rather than restricted formal languages—enabling language-agnostic constraint modeling. We further implement AutoExe, a lightweight LLM-native symbolic execution engine. Experimental results demonstrate that our method substantially improves both analysis accuracy and path-exploration scalability for small-scale LLMs running on consumer-grade hardware, matching the performance of traditional symbolic execution tools. To the best of our knowledge, this is the first work achieving highly accessible and broadly generalizable LLM-native symbolic execution.

Enhancing LLM-based program analysis accuracy and scaleGeneralizing path constraints without formal language translationReducing hardware requirements for symbolic execution tasks

Latest Papers

What's happening recently
View more

This work proposes a novel approach to program analysis and optimization leveraging large language models (LLMs). Addressing the challenge of effectively integrating source code and intermediate representation (IR) information—a limitation in existing methods—it introduces LLMCompiler, pre-trained on IR, and employs a chunked embedding and aggregation strategy to produce unified program-level embeddings. By innovatively unifying the semantics of source code and IR, the method achieves a 1.54% error rate on algorithm classification, representing a 12% improvement over the current state of the art. It also attains competitive accuracy in heterogeneous device mapping, significantly advancing the application of LLMs in program understanding and optimization.

code optimizationintermediate representationLarge Language Models

This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.

abstractionscode performance optimizationhigh-performance computing

This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.

inference pipelinesmicrocontrollerneural inference

This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.

deterministic workflowsLLM agent tracestool-use

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
MR

Manuele Rusci

MSCA Post Doc @ KU Leuven
Embedded SystemsTinyMLLow Power Smart SensorsOn-Device Learning
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
SS

Saransh Sharma

Adobe Research | IIT Kharagpur
Large Language ModelsNatural Language Processing
SE

Sabri Eyuboglu

PhD Student in Computer Science, Stanford University
Machine learning