🤖 AI Summary
This study addresses how to effectively leverage the Geometric and Spectral Alignment (GSA) structure of neural networks to guide model engineering design. To this end, it proposes the CHASE framework, which pioneers the construction of shared attention heads based on geometric alignment and integrates spectral concentration analysis with low-rank subspace extraction to enable inter-layer KV cache sharing and structured pruning. The framework systematically encompasses six application scenarios, including fine-tuning, pruning compensation, and KV cache compression. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines, exhibiting particularly strong performance in model merging, MHA-to-GQA conversion, and KV cache compression tasks. Ultimately, this work establishes a unified and efficient paradigm for multi-task model engineering.
📝 Abstract
Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six applications across model modification, reconfiguration, and compression. CORA, COEC, and CORAM apply GSA to parameter-efficient finetuning, structured-pruning compensation, and model merging. We further develop three new methods. CAGA uses GSA to identify multi-head attention heads that can share a KV representation and constructs the shared key and value heads through geometric alignment and low-rank subspace extraction. SAKV uses GSA to determine which adjacent layers can share a low-rank KV-cache representation and the retained rank for each layer group. CAPS uses GSA spectral structure to group output neurons and selects retained input channels separately for each group. Results from CORA, COEC, and CORAM establish the effectiveness of GSA for adaptation, pruning compensation, and model merging. Experiments on CAGA show that geometric shared-head construction substantially improves MHA-to-GQA conversion, and SAKV and CAPS improve over representative baselines for KV-cache compression and structured pruning. These results show that the structures identified by GSA can be used directly to design methods for a range of model operations.