🤖 AI Summary
Existing evaluation metrics focus solely on final task outcomes or unordered agent routing, failing to capture deviations in the compliance of agent execution trajectories within payment workflows. This work proposes Agentic Success Rate (ASR), a trajectory-level metric that quantifies workflow fidelity by comparing actual and expected agent execution sequences at the transition level, decomposed into transition recall and precision. Leveraging the multi-agent payment system HMASP—integrating LLM-based trajectory analysis, transition-level alignment, prompt engineering, and deterministic routing safeguards—we validate ASR’s effectiveness across 18 large language models and 90,000 task instances. ASR-guided diagnostics yield up to a 93.8-percentage-point improvement in Task Success Rate (TSR) for certain models, with GPT-5.2 achieving perfect ASR, thereby uncovering critical process-skipping issues invisible to conventional metrics.
📝 Abstract
LLM-based multi-agent systems are increasingly deployed for payment workflows, yet prevailing metrics, Task Success Rate (TSR) and Agent Handoff F1-Score (HF1), capture only final outcomes or unordered routing decisions. We introduce the Agentic Success Rate (ASR), a trajectory-fidelity metric that compares observed and expected agent execution sequences at the transition level, decomposing performance into Transition Recall and Transition Precision. Applied to the Hierarchical Multi-Agent System for Payments (HMASP) across 18 LLMs and 90,000 task instances, ASR reveals that 10 of 18 models systematically skip a confirmation checkpoint during payment checkout, a deviation invisible to both TSR and HF1, while 8 models enforce the checkpoint perfectly. Notably, GPT-4.1 exhibits hidden workflow shortcuts despite achieving perfect TSR and HF1, while GPT-5.2 achieves perfect ASR. Prompt refinements and deterministic routing guards guided by ASR diagnostics yield substantial TSR improvements, with gains up to +93.8 percentage points for previously struggling models, demonstrating that trajectory-level evaluation is essential in regulated domains.