Score
Design and implement software artifacts and build processes that embed trained machine learning models into constrained runtime environments, including generating deterministic C modules and compiling inference into application binaries to remove external ML runtime dependencies. Optimize, validate, and integrate the inference code for platform constraints and low-latency execution (for example microsecond-scale inference) while ensuring deterministic behavior.
This paper addresses the prevalent “silent failures” and unpredictable behavior of large language models (LLMs) when auto-generating code for embedded machine learning (ML) workflows. We propose a closed-loop evaluation framework covering data preprocessing, model conversion, and on-device inference code generation. Through multi-model empirical analysis, we introduce the first failure taxonomy for LLM-generated code in embedded ML, identifying systemic fragility arising from prompt format bias, implicit structural assumptions encoded in LLMs, and blind spots in compilation- and runtime-level validation. Key failure patterns include format-misleading parsing errors and “compilable-yet-functionally-broken” runtime errors—both largely undetectable by conventional verification methods. Our findings provide both theoretical foundations and practical guidelines for enhancing the reliability, traceability, and robustness of LLM-driven embedded ML systems.
Deploying large code language models on resource-constrained local devices remains challenging due to hardware limitations, which compromises privacy preservation, inference latency, and offline availability. This work proposes Ditto, a method that co-optimizes model compression and inference program generation to compile large code models into lightweight executables suitable for statically typed languages such as C. Ditto innovatively integrates bounded-error product quantization with quantization-aware inference program synthesis and extends the LLVM compiler to automatically replace general matrix-vector (GEMV) operations with efficient BLAS calls. Evaluated across three prominent code language models, Ditto achieves up to 10.5× speedup and 6.4× reduction in memory footprint, with an average drop of only 0.27% in pass@1 accuracy.
This work addresses the high code complexity and maintenance overhead commonly faced by general-purpose large language model (LLM) inference frameworks due to their need to support diverse models, hardware platforms, and optimization strategies. The authors propose a novel “LLM-as-compiler” paradigm that uniquely integrates explicit knowledge-driven reasoning with multi-agent collaboration, establishing a closed-loop pipeline of generation, verification, and knowledge distillation. In this framework, users specify only high-level inference constraints, and the system—leveraging a Contractual Knowledge Base (CKB) and coordinated multi-agent interaction—automatically synthesizes lightweight, redundancy-free, customized inference engines. Experimental results demonstrate that the approach can successfully generate high-performance, executable inference engines without any reference implementation, thereby validating the feasibility and effectiveness of knowledge-guided automated construction of efficient LLM inference systems.
Large language model inference lacks output determinism due to floating-point non-associativity, dynamic batching, and varying GPU reduction orders. This work proposes a scheduling-based speculative validation mechanism that introduces speculative execution into deterministic inference for the first time. By employing lightweight validate-and-rollback cycles combined with fixed-shape reduction scheduling, the approach incurs overhead only for requests requiring determinism, while remaining compatible with dynamic batching and requiring minimal modification to existing GPU kernels. The method decouples determinism guarantees from low-level implementation details, achieving high throughput and significantly outperforming baseline strategies such as disabling dynamic batching or rewriting kernel functions.
Existing large language models (LLMs) face significant challenges in direct symbolic execution—including low precision, high computational overhead, and strong dependence on large-scale models and high-end hardware—hindering practical deployment. To address these limitations, we propose a novel LLM-driven lightweight symbolic execution paradigm. Our approach employs path-guided task decomposition to decouple complex program analysis into fine-grained, resource-efficient subtasks. We introduce the first path-constraint generalization method based on universal code representations—rather than restricted formal languages—enabling language-agnostic constraint modeling. We further implement AutoExe, a lightweight LLM-native symbolic execution engine. Experimental results demonstrate that our method substantially improves both analysis accuracy and path-exploration scalability for small-scale LLMs running on consumer-grade hardware, matching the performance of traditional symbolic execution tools. To the best of our knowledge, this is the first work achieving highly accessible and broadly generalizable LLM-native symbolic execution.
This work proposes a novel approach to program analysis and optimization leveraging large language models (LLMs). Addressing the challenge of effectively integrating source code and intermediate representation (IR) information—a limitation in existing methods—it introduces LLMCompiler, pre-trained on IR, and employs a chunked embedding and aggregation strategy to produce unified program-level embeddings. By innovatively unifying the semantics of source code and IR, the method achieves a 1.54% error rate on algorithm classification, representing a 12% improvement over the current state of the art. It also attains competitive accuracy in heterogeneous device mapping, significantly advancing the application of LLMs in program understanding and optimization.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.