Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of disentangling execution performance from audit scores in the evaluation of LLM-based financial agents. To isolate the effects of model responses from execution rules, we propose a fixed-tape replay mechanism and construct a multi-defect audit task suite with explicit multi-label prompting to refine answer keys. By integrating synthetic scenario simulation, seed clustering, and Holm correction, this work quantifies the significant negative impact of stressed execution on returns and demonstrates that single-objective recall fails to accurately reflect audit quality. Ultimately, this research delineates the valid boundaries for evaluating LLM financial agents and provides methodological support for reliable scoring frameworks.
📝 Abstract
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase mixes the execution rule with fresh model responses and portfolio feedback: the parsed decision paths agree in only $19.8\%$ of $450$ pairs. Replaying each stored response tape through both execution destinations gives a narrower result. Conditional on those responses, stressed execution changes total return by $-0.0170$ (95\% interval $[-0.0230,-0.0117]$), or $10.4\%$ of the idealized baseline, and ten seed clusters do not resolve the model ranking. Study~B corrects an incomplete answer key and replaces legacy tasks with matched zero-, one-, and two-defect tasks under an explicit multi-label prompt. The drop in target violation recall from one to two defects is positive in five of six combinations of auditor and source (median $0.267$), with three surviving Holm correction. Yet the auditor that includes both target labels most often has micro-precision $0.149$, emits findings on $98/100$ zero-defect tasks, and returns the exact dual-defect set in only $21/100$ cases. Target recall by itself therefore gives a poor account of audit quality on this construction. The studies address different limits: what an execution comparison estimates, and what target recall captures. Together, they show how fixed conditions and diagnostic controls bound the claims a score can support.
Problem

Research questions and friction points this paper is trying to address.

LLM financial agent evaluation
execution performance interpretation
multi-defect auditing
measurement boundaries
audit quality metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Financial Agent Evaluation
Fixed-Tape Replay
Multi-Defect Auditing
Diagnostic Controls
Measurement Boundaries
🔎 Similar Papers
No similar papers found.
W
Weicheng Xue
Virginia Tech