Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
This study investigates whether local reconstruction gain in residual completion serves as a reliable proxy for final model fidelity. Leveraging a frozen Qwen3 backbone, the authors conduct multi-level intervention experiments using both training-free RESA and learnable Top-K+φ methods. The work reveals a non-monotonic relationship between local optimization and global performance: improving local reconstruction accuracy at attention layers does not necessarily enhance output fidelity, and positive local gains can even coincide with degraded KL divergence. These findings challenge the validity assumptions underlying existing evaluation paradigms for sparse attention and clarify the inherent limitations of approximate completion in specific scenarios.