🤖 AI Summary
This study addresses the vulnerability of mathematical agents to silently corrupted tool feedback, which impairs error detection and correction and degrades reliability. To investigate this, we propose a controlled corruption evaluation framework that employs hidden interceptors to simulate tool failures, systematically comparing the robustness of four strategies: no verification, forced in-context reflection, optional new-context verification, and structured verification. Our analysis reveals that verifier availability and verification strategy constitute independent components, demonstrating that forced reflection can fully recover performance, whereas optional verification relies on the model's proactive invocation. Experimental results indicate that without verification, accuracy drops to 72.4%, while both forced reflection and explicit detection followed by restart restore the resolution rate to 100%.
📝 Abstract
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.