Forking: Sudden Overfitting Under Replay
This study addresses the abrupt divergence between training and validation losses at epoch boundaries observed in NanoGPT during data replay. We reveal that this "bifurcation" phenomenon stems from an n-gram memory module that, through repeated updates, amplifies context-specific subspaces while suppressing the probabilities of unseen sequences, thereby triggering sudden overfitting. Through controlled experiments utilizing NanoGPT and DeepSeek-style Engram models, we systematically dissect the underlying n-gram encoding mechanisms and the influence of low-frequency contexts. Our contributions include successfully reproducing and confirming the prevalence of this bifurcation phenomenon across short-budget repetitive scenarios in both supervised fine-tuning and reinforcement learning. This work provides a novel perspective for understanding model generalization failure and highlights potential adverse side effects arising from techniques generated by autonomous AI research agents.