Agent Policy-Value Audit: Separating Transition Composition from Event Selection in Financial LLM Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of end-to-end return testing in evaluating financial LLM agents, which fails to disentangle active selection skills from passive market exposure. To overcome this, we propose a Strategy Value Auditing framework that introduces a novel transfer-matching audit paradigm. By fixing action transition types while randomly reallocating them, this approach effectively isolates selection value from confounding portfolio benchmarks and market drift. Integrating semi-synthetic benchmarks, randomized algorithms, and paired difference tests, the method reduces the false positive rate from 11.6% to 5.3%. Empirical evaluations reveal that existing agents exhibit no significant selection advantage. Consequently, we recommend independently reporting deployment and selection values, establishing a rigorous new standard for financial AI evaluation.
📝 Abstract
Financial LLM agents are often evaluated by comparing their end-to-end returns with those of a baseline and testing the paired difference against zero. This measures whether deploying the agent changes realized performance, but it does not isolate event-selection skill. An agent that frequently changes positions from flat to long can earn a positive paired return from an upward-drifting event pool even when it selects events at random. We propose the Agent Policy-Value Audit, which holds fixed the observed count of each ordered action-change type and randomly reassigns them across eligible events. The average payoff from these reassignments is the composition benchmark; the difference between observed deployment value and this benchmark is selection value. In semi-synthetic benchmarks based on real earnings-event returns, a zero-centered paired test falsely attributes passive exposure to selection skill in $11.6\%$ of no-skill replications, while the transition-matched audit reduces this rate to $5.3\%$. Applied retrospectively to 723 earnings events at 44 U.S. consumer-facing firms, the audit decomposes the agent's gross deployment value of $+15.2$ bps/event into a $+25.8$ composition benchmark and a $-10.6$ selection value. The agent does not detectably outperform matched random assignments. Financial-agent evaluations should report deployment value separately from event-selection value.
Problem

Research questions and friction points this paper is trying to address.

Financial LLM Agents
Event Selection Skill
Agent Evaluation
Composition Benchmark
Policy-Value Audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy-Value Audit
Financial LLM Agents
Event Selection
Transition Composition
Performance Decomposition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mingyang (Alex) Chen
Emory University
Y
Yida (Andrew) Xu
Emory University
H
Huiwen (Aurora) Chen
Emory University
Y
Yiming Lu
Emory University
Wei Jin
Wei Jin
Assistant Professor, Emory University
Graph Neural NetworksMachine LearningAI for ScienceComputational Epidemiology