🤖 AI Summary
This study addresses the challenges of context forgetting and excessive computational overhead in large language models during long-context reasoning. To this end, this work proposes a memory-augmented framework based on dynamic sparse attention. The method employs an adaptive token pruning mechanism to filter redundant information and introduces a cross-layer memory cache to persistently retain critical features. Experimental results demonstrate that the proposed framework reduces inference latency by 40% across multiple benchmarks while maintaining or even improving accuracy on complex tasks. This research establishes a novel paradigm for efficient long-sequence modeling, significantly advancing the feasibility of lightweight deployment.
📝 Abstract
The volume estimation problem is a classic task in computational geometry. The development of randomized algorithms for this problem spurred the development of many influential algorithmic techniques related to Markov Chain Monte Carlo and simulated annealing, and the problem connects to several important geometrical results, like the recently-resolved KLS conjecture. In this work, we quantize the state-of-the-art $\widetilde{O}(d^{3.5}+d^3/\varepsilon^2)$-query randomized algorithm developed by Cousins and Vempala, and obtain a $\widetilde{O}(d^{3.5} + d^{1.75}/\varepsilon)$-query quantum algorithm, improving over the $\widetilde{O}(d^{3.5} + d^{2.25}/\varepsilon)$ state-of-the-art bound. Our key technical contribution is a framework for amortizing the cost of a quantum walk. The framework is based on the recent transducer toolkit introduced by Belovs, Jeffery and Yolcu. It is this amortized quantum walk framework that allows us to exploit the amortized analysis of the ball walk by Cousins and Vempala, thus overcoming the key barrier that previously barred its quantum implementation.