Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems

📅 2026-05-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation metrics focus solely on final task outcomes or unordered agent routing, failing to capture deviations in the compliance of agent execution trajectories within payment workflows. This work proposes Agentic Success Rate (ASR), a trajectory-level metric that quantifies workflow fidelity by comparing actual and expected agent execution sequences at the transition level, decomposed into transition recall and precision. Leveraging the multi-agent payment system HMASP—integrating LLM-based trajectory analysis, transition-level alignment, prompt engineering, and deterministic routing safeguards—we validate ASR’s effectiveness across 18 large language models and 90,000 task instances. ASR-guided diagnostics yield up to a 93.8-percentage-point improvement in Task Success Rate (TSR) for certain models, with GPT-5.2 achieving perfect ASR, thereby uncovering critical process-skipping issues invisible to conventional metrics.
📝 Abstract
LLM-based multi-agent systems are increasingly deployed for payment workflows, yet prevailing metrics, Task Success Rate (TSR) and Agent Handoff F1-Score (HF1), capture only final outcomes or unordered routing decisions. We introduce the Agentic Success Rate (ASR), a trajectory-fidelity metric that compares observed and expected agent execution sequences at the transition level, decomposing performance into Transition Recall and Transition Precision. Applied to the Hierarchical Multi-Agent System for Payments (HMASP) across 18 LLMs and 90,000 task instances, ASR reveals that 10 of 18 models systematically skip a confirmation checkpoint during payment checkout, a deviation invisible to both TSR and HF1, while 8 models enforce the checkpoint perfectly. Notably, GPT-4.1 exhibits hidden workflow shortcuts despite achieving perfect TSR and HF1, while GPT-5.2 achieves perfect ASR. Prompt refinements and deterministic routing guards guided by ASR diagnostics yield substantial TSR improvements, with gains up to +93.8 percentage points for previously struggling models, demonstrating that trajectory-level evaluation is essential in regulated domains.
Problem

Research questions and friction points this paper is trying to address.

workflow fidelity
LLM-based agentic systems
payment workflows
trajectory evaluation
transition-level metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Success Rate
workflow fidelity
trajectory evaluation
multi-agent payment systems
transition-level metrics
🔎 Similar Papers
2024-10-10arXiv.orgCitations: 0
D
Donghao Huang
School of Computing and Information Systems, Singapore Management University, 80 Stamford Rd, Singapore 178902, Singapore; Research and Development, Mastercard, 4250 Fairfax Dr, Arlington, VA 22203, USA
J
Joon Kiat Chua
School of Computing and Information Systems, Singapore Management University, 80 Stamford Rd, Singapore 178902, Singapore
Z
Zhaoxia Wang
School of Computing and Information Systems, Singapore Management University, 80 Stamford Rd, Singapore 178902, Singapore