Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?

๐Ÿ“… 2025-08-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Large language models (LLMs) often produce incorrect final answers in complex reasoning tasks due to undetected errors in intermediate reasoning steps; existing chain-of-thought (CoT) prompting methods lack explicit mechanisms for error identification and correction. To address this, we propose Error-Reflective Prompting (ERP), the first prompting framework that integrates automated error detection, attribution, and correction directly into the CoT process: after generating an initial answer, the model autonomously backtracks through its reasoning trace, pinpoints erroneous steps, constructs an โ€œerror profile,โ€ and regenerates a corrected solution. ERP requires no additional training or fine-tuningโ€”only structured prompting enables self-reflection. Experiments across mathematical reasoning and commonsense question answering demonstrate that ERP significantly improves accuracy and stability while enhancing interpretability and robustness of reasoning traces. ERP thus provides a general, lightweight, prompt-level solution for trustworthy LLM reasoning.

Technology Category

Natural Language Processing: Prompt Engineering / PromptingKnowledge Representation and Reasoning: Common-Sense ReasoningReasoning under Uncertainty: Uncertainty Representations

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
๐Ÿ“ Abstract
Prompting methods for language models, such as Chain-of-thought (CoT), present intuitive step-by-step processes for problem solving. These methodologies aim to equip models with a better understanding of the correct procedures for addressing a given task. Despite these advancements, CoT lacks the ability of reflection and error correction, potentially causing a model to perpetuate mistakes and errors. Therefore, inspired by the human ability for said tasks, we propose Error Reflection Prompting (ERP) to further enhance reasoning in language models. Building upon CoT, ERP is a method comprised of an incorrect answer, error recognition, and a correct answer. This process enables the model to recognize types of errors and the steps that lead to incorrect answers, allowing the model to better discern which steps to avoid and which to take. The model is able to generate the error outlines itself with automated ERP generation, allowing for error recognition and correction to be integrated into the reasoning chain and produce scalability and reliability in the process. The results demonstrate that ERP serves as a versatile supplement to conventional CoT, ultimately contributing to more robust and capable reasoning abilities along with increased interpretability in how models ultimately reach their errors.
Problem

Research questions and friction points this paper is trying to address.

Enhancing error recognition and correction in language models
Addressing Chain-of-Thought's inability to reflect on mistakes
Improving reasoning reliability through error analysis steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Error Reflection Prompting enhances reasoning through error recognition
Automated ERP generation integrates error correction into reasoning chains
ERP supplements Chain-of-thought with scalable error analysis