MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

πŸ“… 2026-07-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
δΈΊι™δ½Žε€§εž‹θ―­θ¨€ζ¨‘εž‹ζŽ¨η†ζˆζœ¬οΌŒζε‡ΊCacheSpecζ‘†ζžΆοΌŒεˆ©η”¨ε°εž‹ζ¨‘εž‹θΏ›θ‘Œθ―­δΉ‰ε˜ι‡ζε–ε’ŒζŽ¨ζ΅‹θ‰η¨Ώη”ŸζˆοΌŒζι«˜ηΌ“ε­˜ε€η”¨ηŽ‡ε’Œε€„η†ι€ŸεΊ¦γ€‚
πŸ“ Abstract
Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applications often incur high inference cost. We present MiniCache, a reusable program caching framework that transforms Program-of-Thought (PoT) programs into parameterized cache objects, enabling reusable computation across structurally similar requests. MiniCache reuses the same small model for semantic variable extraction on cache-hit requests and speculative drafting during target-LLM generation, reducing expensive target-LLM invocations while preserving task quality. Experiments on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA demonstrate that MiniCache improves the trade-off between inference latency, cache reuse, and accuracy, achieving up to 3.1x lower latency and 2.8x higher throughput under parallel serving. These results show that small models are most effective not as replacements for large models, but as lightweight interface models that enable reliable and efficient reusable program caching.
Problem

Research questions and friction points this paper is trying to address.

large language models
inference cost
program-level caching
small models
reusable computation logic
Innovation

Methods, ideas, or system contributions that make the work stand out.

program-level caching
small models
inference optimization
reusable computation logic
semantic variable extraction
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
J
Jingquan Chen
University of Electronic Science and Technology of China
J
Jie Feng
Zhongguancun Academy
Jinghua Piao
Jinghua Piao
Tsinghua University
S
Shaogang Hu
University of Electronic Science and Technology of China
Y
Yong Li
Department of Electronic Engineering, BNRist, Tsinghua University