🤖 AI Summary
This study addresses the challenge of automatically coordinating knowledge domain sharing and specialization during large model training. To this end, it proposes CORTEX, a framework that pioneers the internal modularization of dense models based on learning dynamics. Methodologically, CORTEX leverages domain-conditioned gradients and cross-domain similarity to group matrix parameters, automatically partitioning them via gradient similarity to balance shared and specialized capabilities. Furthermore, it introduces selective lesion scoring and module-domain mutual information to quantify alignment effectiveness. Experiments demonstrate that this approach achieves optimal synthetic domain matching across models of varying scales while significantly reducing perplexity. Simultaneously, it preserves performance on real-world domains and yields interpretable parameter modules.
📝 Abstract
Large language models are trained on heterogeneous data mixtures, where different knowledge domains require both shared knowledge and specialization. Existing modular approaches typically impose explicit components or discover modules through interpretability analysis after training. In this work, we propose CORTEX, a learning dynamics-inspired framework that learns internal modularization within dense language models. CORTEX partitions trainable matrices into parameter groups and learns module assignments from domain-conditioned gradient and cross-domain gradient similarity. We introduce the selective lesion score and module-domain mutual information to characterize the target-domain lesion effects and alignment, and analyze how module assignment affects the trade-off between assignment bias and update magnitude. Experiments with 160M, Qwen3-8B, and Qwen3-32B backbone models show that CORTEX achieves the highest synthetic-domain exact match and largest average perplexity reduction, while remaining competitive on real-domain evaluations and forming identifiable modules.