🤖 AI Summary
This study addresses the evaluation bias in sparse accelerator design caused by reliance on analytical models by proposing a large language model (LLM)-driven closed-loop hardware-software co-optimization framework. Leveraging the Model Context Protocol (MCP), the framework enables real-time editing of RTL, memory configurations, and kernel scheduling. It integrates legality checks, cycle-accurate simulation, and bit-level verification, while employing auxiliary models to automatically repair failed designs for iterative optimization guided by hardware measurement feedback. Experimental evaluations on the Gemmini architecture demonstrate that, compared to the baseline, the proposed approach reduces loop iterations by 2.1× and off-chip traffic by 9.8×, decreases area by 22.8%, improves energy efficiency by 5.61×, and lowers the energy-delay product by 11.8×.
📝 Abstract
Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a $512 \times 512$ GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.