CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering
This study addresses how to effectively leverage the Geometric and Spectral Alignment (GSA) structure of neural networks to guide model engineering design. To this end, it proposes the CHASE framework, which pioneers the construction of shared attention heads based on geometric alignment and integrates spectral concentration analysis with low-rank subspace extraction to enable inter-layer KV cache sharing and structured pruning. The framework systematically encompasses six application scenarios, including fine-tuning, pruning compensation, and KV cache compression. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines, exhibiting particularly strong performance in model merging, MHA-to-GQA conversion, and KV cache compression tasks. Ultimately, this work establishes a unified and efficient paradigm for multi-task model engineering.