When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in large language models (LLMs) performing intrinsic self-correction without external evidence, where rectifying errors risks inadvertently altering correct responses. Rather than treating self-correction as a uniformly beneficial process, this work reconceptualizes it as a revision strategy that necessitates jointly evaluating correction gains against the risk of introducing errors. Through an empirical analysis of 29 open-source LLMs, the authors employ correctness state transition tracking and selective invocation mechanisms to optimize decision-making. The findings reveal underlying behavioral discrepancies obscured by aggregated accuracy metrics and identify specific scenarios where learned gating strategies outperform unconditional correction. Ultimately, this research establishes a more rigorous evaluation framework and offers practical guidance for advancing LLM self-correction capabilities.
📝 Abstract
Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence. A second pass can recover mistakes, but it can also overturn answers that were already correct. We study this trade-off across 29 open-weight LLMs on BoolQ, GSM8K, and Corr2Cause by tracking correctness transitions between initial and revised answers. Aggregate accuracy can conceal substantially different revision behavior: for example, Llama-3.1-8B improves by 25.5 percentage points on GSM8K, while refinement changes 19.1% of initially correct answers into wrong ones. A controlled BoolQ study further shows that refinement prompts shift the balance between recovery and harm. We then compare three runtime choices: keeping the initial answer, always accepting the revision, and selectively invoking revision using signals available after the initial response. The comparison identifies settings where learned gating is useful and others where a simpler unconditional policy performs better. These results suggest treating intrinsic self-correction as a revision policy rather than as a uniformly beneficial second pass, and evaluating it through both the corrections it recovers and the errors it introduces.
Problem

Research questions and friction points this paper is trying to address.

Intrinsic self-correction
Large language models
Revision policy
Risk-aware evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intrinsic Self-Correction
Revision Policy
Risk-Aware Evaluation
Learned Gating
Correctness Transitions
🔎 Similar Papers
No similar papers found.