🤖 AI Summary
This work addresses the tendency of large language model (LLM) agents to discard previously validated successful corrections during task repair, leading to redundant exploration. To mitigate this, the authors propose MERIT, a training-free agent that introduces, for the first time, a causality-aware, type-guided bipolar memory mechanism to explicitly distinguish between successful and failed repair trajectories. Built upon a frozen LLM, MERIT enables lightweight cross-query experience reuse through a deterministic failure classifier and a hybrid lexical-dense retriever. Evaluated on the Spider and BIRD datasets, MERIT improves execution accuracy from 66.34% to 69.79% and from 47.35% to 48.44%, respectively, demonstrating the efficacy of its causal memory architecture.
📝 Abstract
LLM agents that repair failures often discard successful corrections, forcing later episodes to rediscover similar solutions. We study whether finalized repair outcomes can improve subsequent Text-to-SQL episodes without parameter updates. We introduce MERIT, a training-free agent that maintains an online dual-polarity memory of oracle-verified corrections and observed unsuccessful directions. Under oracle-assisted benchmark feedback, only memories from earlier finalized episodes are eligible for retrieval. A deterministic classifier assigns a coarse failure type, which conditions a hybrid lexical-dense retriever before the frozen model generates each revision. Using Qwen2.5-7B-Instruct with identical initial predictions and repair budgets, MERIT improves execution accuracy over stateless iterative repair from \(66.34\%\) to \(69.79\%\) on Spider and from \(47.35\%\) to \(48.44\%\) on BIRD. Paired analyses provide clear evidence for the Spider gain but weaker evidence on BIRD. MERIT is not reliably separated from untyped dynamic retrieval on either benchmark, while Reflexion-style memory reaches \(51.24\%\) on BIRD at substantially higher inference cost. Ablations show that negative memory contributes modestly, the value of type conditioning and lexical--dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit. These results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable.