Forking: Sudden Overfitting Under Replay

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the abrupt divergence between training and validation losses at epoch boundaries observed in NanoGPT during data replay. We reveal that this "bifurcation" phenomenon stems from an n-gram memory module that, through repeated updates, amplifies context-specific subspaces while suppressing the probabilities of unseen sequences, thereby triggering sudden overfitting. Through controlled experiments utilizing NanoGPT and DeepSeek-style Engram models, we systematically dissect the underlying n-gram encoding mechanisms and the influence of low-frequency contexts. Our contributions include successfully reproducing and confirming the prevalence of this bifurcation phenomenon across short-budget repetitive scenarios in both supervised fine-tuning and reinforcement learning. This work provides a novel perspective for understanding model generalization failure and highlights potential adverse side effects arising from techniques generated by autonomous AI research agents.
📝 Abstract
This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily crowded tables suppress it. We also observe forking in short-budget, heavily repeated SFT and RL-like regimes. The contributions of this paper are twofold: (1) Forking reveals yet another curious phenomenon in deep learning, in addition to grokking and double descent. (2) Forking is an unexpected and unpleasant by-product of tricks proposed by autoresearch agents. While these agents produce an enormous number of results that seem useful, we should always be careful with their results.
Problem

Research questions and friction points this paper is trying to address.

Forking
Overfitting
Data Replay
NanoGPT
Autoresearch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Forking
Data Replay
N-gram Memory
Autoresearch Agents
Overfitting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shanbin Yu
MetaCircle, Tsinghua University
Shaoyang Guo
Shaoyang Guo
Peking University
PhysicsAI
H
Haoran Zhao
MetaCircle, Peking University
Danni Yu
Danni Yu
MetaCircle, Tsinghua University
Z
Ziming Liu
MetaCircle, Shanghai Qizhi Institute, Tsinghua University