🤖 AI Summary
This study addresses the challenge of rapidly distinguishing individual sample contributions and the lack of controllability in behavioral interventions during large model training. To this end, it introduces mutual information to quantify intra-batch interference and proposes a novel Behavior Gradient Uniqueness (BGU) metric grounded in geometric interpretation, alongside the BS-Ghost algorithm. These innovations enable gradient-free batch-space shared computation and real-time weight intervention. Consequently, this work presents a low-overhead framework for training-time attribution and dynamic control. Evaluated on Qwen2.5-7B, the proposed approach incurs only an 8% increase in training time while accurately identifying causal data samples and steering model evolution toward target behaviors, thereby offering an efficient and controllable paradigm for optimizing large model training.
📝 Abstract
Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.