FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing agent evaluation benchmarks struggle to assess capabilities in experience transfer and self-evolution across continuous professional tasks, and lack comprehensive evaluation of open-ended outputs and multi-dimensional compliance within financial workflows. This work proposes the first longitudinal benchmark tailored to the financial domain, encompassing six areas, 20 scenarios, and 120 real-world tasks. Integrating domain-specific procedural constraints and human review, it unifies financial workflows, open deliverables, and multi-faceted assessment into a quantifiable framework for measuring agent self-evolution efficacy, introducing dual metrics for task quality and compliance. Evaluating four self-evolution architectures atop the Qwen3.7-Max backbone, experiments show that Letta achieves the highest post-evolution score (91.65) and lowest compliance violation rate (0.09 per task), while Codex exhibits the largest performance gain (+19.37). Overall, self-evolution strategies significantly enhance performance (+9.33 to +19.37 points) and reduce compliance risks (−0.12 to −0.44 violations per task).
📝 Abstract
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold's self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.
Problem

Research questions and friction points this paper is trying to address.

self-evolving agents
longitudinal benchmark
professional workflows
financial compliance
task evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-evolving agents
longitudinal benchmark
financial workflows
compliance-aware evaluation
professional task evolution
🔎 Similar Papers
B
Bo Deng
Beihang University; Qwen DianJin Team, Alibaba Cloud Computing
K
Kang Zhou
Qwen DianJin Team, Alibaba Cloud Computing
Lifan Guo
Lifan Guo
Researcher Drexel University
Machine Learning
Chongyang Tao
Chongyang Tao
Associate Professor of Computer Science, Beihang University
Natural Language ProcessingDialogue SystemsInformation RetrievalData Intelligence
X
Xuanren Chen
Beihang University
C
Chenggang Xie
Beihang University
R
Renzhao Liang
Beihang University
F
Feng Chen
Qwen DianJin Team, Alibaba Cloud Computing
C
Chi Zhang
Qwen DianJin Team, Alibaba Cloud Computing