FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决金融搜索中信息检索、来源选择等问题,引入FinFIRST基准,通过原子评分标准评估答案及支撑证据。
📝 Abstract
Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominantly evaluate only the final answer, making it difficult to localize errors or assess whether an answer is well-founded. To address this gap, we introduce FinFIRST (Financial Information Retrieval, Sourcing and Traceability), the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics. FinFIRST comprises 123 expert-authored tasks spanning a graduated difficulty spectrum, constructed from aggregate patterns of real-world financial scenarios through an 18-field taxonomy, a six-axis coverage blueprint, a registry of 138 financial sources, contributions from over 50 finance experts, and a six-stage quality-control pipeline. Each task is accompanied by an evidence-grounded reference package decomposed into atomic criteria across three dimensions: raw-information acquisition, source verification, and computation and answer formation. We evaluate 15 model configurations under a unified tool setting. Claude-Opus-5 achieves the highest atomic score of 87.59%, while GPT-5.6-Sol attains the highest strict pass rate of 71.54%. Computation and answer formation consistently lag behind raw-information acquisition across systems. FinFIRST retains final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.
Problem

Research questions and friction points this paper is trying to address.

Financial Search
Benchmarking
Information Retrieval
Sourcing
Traceability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Financial Information Retrieval
Sourcing and Traceability
Atomic Rubrics
Evidence-Grounded Evaluation
Quality-Control Pipeline
💼 Related Jobs
No related jobs found.
Wenqing Wang
Wenqing Wang
Postdoctoral researcher, Iowa State University
Dynamic ModelingHierarchical ControlModel Predictive ControlStochastic Control
H
Haitao Xiang
Ling Team, Inclusion AI
Xinyi Zhao
Xinyi Zhao
Columbia university
Data ScienceData Visualization
M
Mingming Yin
Ling Team, Inclusion AI
Ying Zhong
Ying Zhong
Associate Professor of Management Science, University of Electronic Science and Technology of China
Simulation OptimizationMulti-Armed Bandit ProblemsParallel Computing
Z
Zhaoxin Huan
Ling Team, Inclusion AI
Q
Qiheng Zhou
China International Capital Corporation Limited
J
Jin Zhu
Ling Team, Inclusion AI
X
Xiaolu Zhang
Ling Team, Inclusion AI
Shi Chang
Shi Chang
Cornell University
active perceptionsensor motion planning
J
Jun Zhou
Ling Team, Inclusion AI