Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that self-play struggles to generate verifiable cases in evidence-identifiable tasks by proposing a counterfactual self-evolution framework. The method employs instruction tuning to train a Proposer for generating causally interpretable evidence edits, which are optimized via a feedback reward mechanism and organized into a contextual memory bank. Subsequently, counterfactual contexts are leveraged to rectify errors and enhance decision confidence, enabling dynamic reasoning adaptation for a frozen Solver without weight updates. Experimental results demonstrate that the proposed approach achieves superior performance across clinical diagnosis, fact verification, and business reasoning tasks, significantly outperforming multiple state-of-the-art models.
📝 Abstract
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
Problem

Research questions and friction points this paper is trying to address.

counterfactual reasoning
self-play
evidence-grounded reasoning
self-evolving agents
verifiable tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Self-Evolution
Self-Play Proposer-Solver
Evidence-Grounded Reasoning
Instruction Tuning
Contextual Memory
🔎 Similar Papers
No similar papers found.