Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the evaluation challenges of financial AI agents regarding quantitative execution, resource utilization, and auditable outputs by proposing FinSkillBench. This benchmark introduces a novel task design incorporating hidden ground truths for financial scenarios, combining deterministic verifiers, paired analysis, and clustered bootstrap statistics to systematically deconstruct the "skill premium" and examine how skill bundles affect workflows. Empirical findings reveal that curation skills improve average model scores by 16.2 points. Furthermore, tool empowerment significantly outperforms document-based guidance (+19.5 vs. +5.6 points) and dominates in numerically intensive tasks, uncovering the systematic dependencies underlying performance gains.
📝 Abstract
Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation suite of 2,603 point in time episodes across 12 subtasks in portfolio construction, risk management, and fundamental analysis, with hidden regenerable ground truth and task specific deterministic verifiers. Executing 17,820 episodes across 9 models and 3 resource conditions, the paired analysis across 8 models shows that curated skill packages raise mean scores by +16.2 points (0.366 to 0.528), whereas skills generated within a single episode add only +0.5 points while consuming more tokens and turns. We then decompose the curated premium by granting human authored procedural documents and executable domain tools separately: documents alone add +5.6 points, tools alone add +19.5 points, and their combination is subadditive. The premium is strongly workflow dependent: executable tools dominate numerically intensive workflows, documentation matters more when procedural or output schema guidance is the bottleneck, and interpretive tasks benefit from both. The effects are sign stable across 10 scoring variants and cluster bootstrap analyses, and an independently implemented second harness reproduces the directional pattern while showing that effect magnitudes depend on how tools and data are exposed. Overall, a measured "skill premium" is a property of the full model, resource, and harness system rather than of the underlying model alone.
Problem

Research questions and friction points this paper is trying to address.

Financial AI Agents
Skill Premium
Workflow Evaluation
Verifiable Execution
Tool Augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

FinSkillBench
Skill Premium Decomposition
Deterministic Verifiers
Financial AI Agents
Tool-Documentation Subadditivity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jermyn Zhen Yong Bek
Independent Researcher, Singapore, Singapore
Z
Zhuang Qiang Bok
Deep Insight Labs, Singapore, Singapore
Zhongtian Sun
Zhongtian Sun
University of Cambridge
Artificial IntelligenceRepresentation LearningGeometric Machine LearningNeuroscience