ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems
This study addresses the challenges of tracing early-step errors in LLM agents and the limited attribution capability of existing hidden-state representations by proposing ReCast. This method extracts complementary pattern and bias features through layer selection and feature engineering, and trains an encoder via contrastive learning with a ranking objective to transform frozen LLM hidden states into step-level representations tailored for root cause localization. The main contributions include the ReCast framework and the release of the ReCast-2K dataset. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on the Hit@1 metric across four benchmarks, outperforming the strongest baseline by 5.65 and 9.19 percentage points, respectively.