From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of knowledge transfer for large language models (LLMs) on data-scarce neural processing units (NPUs), as well as the uneven credit assignment and high retrieval overhead inherent in existing memory agents. To this end, we propose SAGE, a framework that optimizes credit assignment via Adoption-Tracking Utility (ATU) estimation and explicit adoption records. By employing utility-gated integration to filter high-value experiences and consolidate them into persistent rules, SAGE enables low-overhead cross-task reuse of hardware-specific knowledge. Experimental results demonstrate that SAGE achieves an operator execution rate of 95.5%, with 86.9% of its operators outperforming torch_npu in performance, while delivering up to a 43.99× speedup for sparse Flash attention.
📝 Abstract
High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming but depend on scarce expert data and substantial training compute. Memory-learning agents instead adapt through external memory, but their uniform credit assignment gives adopted and unused experiences the same reward target, potentially biasing subsequent retrieval rankings. Moreover, when learned values guide only retrieval, high-value experiences that generalize across operators must be retrieved repeatedly rather than retained in context, thereby increasing retrieval overhead and weakening cross-task guidance. We therefore present SAGE, a persistent self-improving agent for NPU kernel synthesis. Adoption-Traced Utility estimation (ATU) combines explicit adoption records with kernel evaluation outcomes for adoption-aware credit assignment. Utility-Gated Consolidation (UGC) uses positive utility and repeated adoption across operators to select and abstract reusable rules into a bounded resident context. On NPUKernelBench, SAGE achieves a 95.5% execution rate versus 84.1% for the strongest controlled baseline, with 86.9% of solved operators outperforming torch_npu. With GLM-5.3, SAGE achieves a 43.99x speedup over the torch_npu reference on sparse flash attention. These results show that adoption-aware credit assignment and selective consolidation enable agents to accumulate and reuse hardware-specific knowledge across tasks.
Problem

Research questions and friction points this paper is trying to address.

NPU kernel synthesis
data-scarce domain-specific architectures
memory-learning agents
credit assignment
knowledge transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adoption-Aware Credit Assignment
Utility-Gated Consolidation
Memory Learning
NPU Kernel Synthesis
Self-Improving Agent
L
Longxiao Fan
Fudan University
T
Tao Zhang
University of Science and Technology of China
H
Han Yan
Fudan University
J
Jiajun Li
Huawei Technologies Ltd.
Mingcong Song
Mingcong Song
University of Florida
computer architecture/systemmachine learning
G
Guoping Long
Huawei Technologies Ltd.
H
Hongjie Si
Huawei Technologies Ltd.
Weiwei Sun
Weiwei Sun
Fudan University
computer science