FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost, poor scalability, and limited institutional adaptability of manually crafted rubrics for evaluating financial research agents by proposing the first expert-guided framework for automated rubric generation. The method employs a collaborative multi-agent architecture comprising writer and reviewer agents, integrated with code execution verification, long-horizon iterative reasoning, and human-in-the-loop mechanisms to enable automatic generation, auditing, and refinement of rubrics while supporting task bank reusability. Experimental results demonstrate that the generated rubrics achieve strong alignment with human scoring across three benchmarks and are preferred in blind evaluations conducted by internal analysts. Furthermore, this work introduces FinAutoRubric, a dataset comprising 100 queries designed to facilitate future research in automated evaluation for financial AI systems.
📝 Abstract
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts'key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.
Problem

Research questions and friction points this paper is trying to address.

financial research agents
rubric generation
evaluation benchmarks
expert standards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Rubric Generation
Expert-Guided Evaluation
Multi-Agent Collaboration
Financial Research Agents
Long-Horizon Loops
💼 Related Jobs
No related jobs found.
Hoyoung Lee
Hoyoung Lee
Ulsan National Institute of Science and Technology (UNIST)
AI in FinanceFinancial NLPTrustworthy AILarge Language Models
S
Suyeol Yun
LinqAlpha
J
Jack Haverty
LinqAlpha
Y
Yunju Cho
LinqAlpha
M
Meesong Kim
LinqAlpha
D
Daekyung Park
LinqAlpha
S
Sumin Kim
LinqAlpha
Jihoon Kwon
Jihoon Kwon
Seoul National University / Hanwha systems
Radar signal processingRadar machine learningTracking filterMicrowave applications
J
Jasmine Jia Geng
MassMutual Life Insurance
A
Andrew Chin
AllianceBernstein
Y
Yin Luo
Wolfe Research
E
Edward Tong
Google
Y
Yu Yu
BlackRock
Z
Zach Golkhou
J.P. Morgan Chase
M
Minkyu Kim
State Street Corporation
Igor Halperin
Igor Halperin
Fidelity
FinanceAImachine learningphysics
Young Cha
Young Cha
McLean Hospital/Harvard Medical School
Metabolic reprogramming
Alejandro Lopez-Lira
Alejandro Lopez-Lira
Assistant Professor of Finance, University of Florida
FintechMachine LearningAsset PricingMacro FinancePrivate Equity
C
Chanyeol Choi
LinqAlpha
Y
Yongjae Lee
LinqAlpha