🤖 AI Summary
This study investigates whether local reconstruction gain in residual completion serves as a reliable proxy for final model fidelity. Leveraging a frozen Qwen3 backbone, the authors conduct multi-level intervention experiments using both training-free RESA and learnable Top-K+φ methods. The work reveals a non-monotonic relationship between local optimization and global performance: improving local reconstruction accuracy at attention layers does not necessarily enhance output fidelity, and positive local gains can even coincide with degraded KL divergence. These findings challenge the validity assumptions underlying existing evaluation paradigms for sparse attention and clarify the inherent limitations of approximate completion in specific scenarios.
📝 Abstract
Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the final model output. We study training-free RESA and learned Top-K+$\phi$ with frozen backbone language models. A prespecified single-layer screen yields two Qwen3-0.6B/Multi-LexSum interventions for which direct-runtime measurements show positive prespecified request-aggregate local reconstruction gain but worse final KL fidelity than the corresponding all-abstain Exact Top-K baseline on both discovery and prompt-token-disjoint holdout requests. Exact restoration at the same layer instead improves final fidelity, showing that the reversal is specific to approximate completion in these cases. In complementary multi-layer experiments, a task-independent local diagnostic often repairs the tested completion estimators, although the repaired models do not consistently outperform Exact Top-K. Together, these results show that better local reconstruction need not translate into better final-model fidelity.