SPECTRUM: Proximal Spectral Modulation for Looped Self-Distillation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the collapse of code generation diversity as accuracy improves during iterative self-distillation. We propose a recursive self-distillation framework with a fixed information budget, enabling self-evolution without external evaluation. To maintain solution space richness during inference, we introduce Full-Rank Proximal Spectral Modulation, which re-estimates loss-sensitive key-value geometry via reference anchors, establishing a recursive refinement paradigm explicitly aimed at preserving correct solutions. Experiments demonstrate that our approach retains 89.9% of the initial AST diversity on MBPP, significantly outperforming baseline models. Furthermore, evaluations on benchmarks such as HumanEval+ confirm the transferability of these diversity advantages.
📝 Abstract
A model that learns from its own outputs inherits more than their correctness: it inherits which solutions it produces. We formulate Looped Self-Distillation, a self-evolution framework for code generation in which a model repeatedly generates and learns from its own raw outputs, under a fixed information budget, without ongoing external assessment or test-based selection of the generated samples. We identify a consequential separation: correctness can improve while the breadth of correct implementations contracts. We introduce SPECTRUM, which re-estimates loss-sensitive key/value geometry from a fixed reference anchor at each round and converts it into full-rank proximal spectral modulation. All generated completions train a single student, whose subsequent inference requires no intervention. After five rounds of experiments on MBPP, SPECTRUM retains 89.9% of the initial model's 64-sample correct AST richness, compared with 66.4% for Vanilla self-distillation and 65.5% for a subspace-projection control. The advantage persists at matched correct-sample counts. Without further training or recalibration, the resulting student also achieves higher matched-correct richness than Vanilla SD on HumanEval+ and APPS Intro, demonstrating transfer of the diversity benefit. These findings establish correct-solution retention as a complementary objective of recursive self-improvement (RSI) and show that generation-time intervention can improve the solution repertoire retained by subsequent students.
Problem

Research questions and friction points this paper is trying to address.

Looped Self-Distillation
Code Generation
Solution Diversity
Recursive Self-Improvement
Self-Evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Self-Distillation
Proximal Spectral Modulation
Code Generation
Recursive Self-Improvement
Diversity Retention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yunbo Long
Yunbo Long
PhD Student, University of Cambridge
Deep LearningGenerative ModelsSynthetic Data
W
WenJie Chen
Fudan University
J
Jiaquan Zhang
Fudan University
G
Guangya Hao
University of Cambridge
Z
Zihang Zeng
Fudan University
P
Pengze Li
Fudan University
X
Xi Chen
Fudan University