GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

πŸ“… 2026-07-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current evaluations of financial models rely on a single ground-truth answer, overlooking the legitimate disagreements inherent among professional analysts and thereby risking misjudgment of AI systems. This work proposes GAUGE, a novel benchmark grounded in 1,001 real analyst workbooks and 196 distinct tasks, which introduces the first evaluation framework anchored in collective professional practice. GAUGE employs a three-tiered practice envelope, 56 auditable dimensions, eight validity gates, and a failure-aware scoring mechanism to enable multidimensional, structured assessment of valuation models. Experimental results show that under the Ο†β‚€ metric, senior analysts achieve an average score of 88.3, while the best AI agent scores 53.4β€”surpassing student-level performance yet significantly lagging behind humans on judgment-intensive tasks. This gap highlights a core limitation of current AI: proficiency in modeling but deficiency in value-based reasoning.
πŸ“ Abstract
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $Ο†_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.
Problem

Research questions and friction points this paper is trying to address.

financial valuation
model evaluation
benchmarking
analyst disagreement
golden answer
Innovation

Methods, ideas, or system contributions that make the work stand out.

GAUGE
financial valuation benchmark
agent-built models
observed-practice envelope
multi-faceted evaluation