🤖 AI Summary
This work addresses the suboptimality in traditional compression-based scoring for model pruning, which, despite being reproducible, suffers from inter-group robustness bias. The authors propose a group-aware pruning method for large language models that prioritizes worst-group performance guarantees. By modeling scoring limitations through an information boundary framework, they introduce a dual-world construction and observational fiber radius to uncover uncoupled factors. To recover degraded ranking fidelity, they develop group-resolved diagonal moment estimation and a depth-allocation strategy. Their approach integrates group-local moments, pooled moment analysis, reference-path curvature modeling, and router trajectory prediction, evaluated via KL divergence and perplexity. Experiments on three dense large language models demonstrate 12.6–20.9% reduction in worst-group perplexity inflation, 2.7–8.0% improvement in endpoint selection over baselines, and 13.7% and 7.2% reductions in worst-group KL divergence using single-layer limited-menu decisions.
📝 Abstract
A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gain. Its selected endpoint was 6.0% and 7.7% worse than two controls. We model the gap through information interfaces that delimit which distinctions each statistic supports. For equal-weight groups, a conic law gives the exact pooling price for positive linear fixed-candidate damage, including diagonal and full PSD second moments. Three two-world constructions and an exact observation-fiber radius characterize what pooled moments, group-local moments, and reference-path curvature leave unresolved. A group-resolved diagonal recovers broad damage order (Spearman 0.9239) while fine order remains weak. Relative to balanced uniform allocation, a coarse depth allocation cuts worst-group perplexity inflation by 12.6--20.9% across three dense LLMs. Model-specific complete-mask endpoint selection improves over those references by 2.7--8.0%. In OLMoE, router traces predict singleton direction (114/192 versus 81/192 under the strongest relabeling). Finite-menu decisions on one layer yield held-out worst-group KL reductions of 13.7% and 7.2%. Local measurements construct candidates. Selection is licensed by complete candidate endpoints or a validated uniform guarantee, with uncertainty calibrated to every comparison.