Model-Aware Data Selection from In-and-Out Information Interplay

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of efficiently selecting data containing unknown knowledge during the post-training of large language models. By revealing complementary patterns between weight and hidden-layer ranks alongside an internal-external information interaction mechanism, this work proposes CAP, a model-aware data selection method. CAP quantifies data accessibility and evaluates its informational value by computing the divergence in inter-layer representational differences between generated and reference responses, thereby enabling precise data selection. Moving beyond traditional heuristic approaches, CAP achieves an average improvement of 35.4% over baselines on mathematical and coding tasks. Notably, it matches full-data training performance using only 10% of the data while demonstrating strong transferability to multimodal settings.
📝 Abstract
LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and"know what they know."We observe an interesting rank equilibrium between knowledge stored in the weights and the data stream passing through the model. Across all model layers, we find that the hidden states (data stream) follow a U-shaped pattern, showing substantial compression in early layers and a steep rise during the late-layer decoding phase. In contrast, the weight rank follows an inverted U-shaped pattern, with very low rank in the early and late layers and high rank in the middle. We interpret this as an in-and-out information interplay: intermediate activations do not need to carry content that the weights can supply later, so they primarily preserve what the weights cannot provide. Motivated by this observation, we propose a model-aware data selection method, CAP (Counterfactual Assimilation Profile), which can determine whether a data candidate contains information accessible to the current model by utilizing the divergence gap in early- and late-layer representations between model-generated and reference responses. Across math, code, and science domains, CAP delivers 35.4% greater average improvement over the base model than the strongest baseline under different selection budgets. With only 10% of the data pool, CAP surpasses or matches full-pool training on math and science. We further show that CAP transfers to multimodal data selection and is robust to response horizon and noise.
Problem

Research questions and friction points this paper is trying to address.

data selection
large language models
post-training
model-aware
knowledge assimilation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-Aware Data Selection
Counterfactual Assimilation Profile
Rank Equilibrium
Information Interplay
Large Language Models
🔎 Similar Papers
No similar papers found.