How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of improper weight allocation in self-correction learning for coding agents, caused by the conflation of evidence support and quality. To this end, we propose Effective Evidence Self-Distillation (EESD), which innovatively decouples execution relevance from evidence quantity. Specifically, EESD leverages Dirichlet posterior modeling to generate uncertainty penalty weights and optimizes the correction learning process under a KL divergence constraint, achieving precise probability estimation and weight control via adaptive pseudocounts. Experimental results demonstrate that our method reduces negative log-likelihood (NLL) by 55%–59% and improves Pass@1 to 20.4% on the DeepSeek/CodeARC benchmark, significantly outperforming fixed-quality baselines.
📝 Abstract
Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
Problem

Research questions and friction points this paper is trying to address.

coding agent
self-correction
evidence weighting
self-distillation
execution feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Dirichlet Evidence
Coding Agent
Self-Correction
Uncertainty Penalization
🔎 Similar Papers
No similar papers found.
Yunbo Long
Yunbo Long
PhD Student, University of Cambridge
Deep LearningGenerative ModelsSynthetic Data
G
Guangya Hao
University of Cambridge
Y
Yuhan Liu
University of Cambridge
Y
Yiting Duan
Western Sydney University
L
Longyan Tan
University of Cambridge
Y
Yunchen Long
Guangdong University of Technology
H
Hao Wu
Western Sydney University