🤖 AI Summary
This study addresses the limitation of existing language agents in effectively correlating historical actions with environmental feedback, which prevents interaction experience from translating into decision-making advantages. To overcome this, we propose an action calibration mechanism that reconstructs “action-observation” causal links through explicit outcome annotation. Furthermore, we design a learnable selective experience recording calibration module that integrates prompt engineering with reinforcement strategies to optimize subsequent decision logic. Our work demonstrates that simple outcome annotation significantly enhances agent performance. Experimental results indicate that the proposed method substantially improves task success rates while reducing ineffective repetitive actions, thereby enabling historical interactions to be genuinely transformed into actionable experience.
📝 Abstract
Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.