SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

πŸ“… 2025-04-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Large language models (LLMs) lack reliable, domain-specific evaluation benchmarks for real-world financial analysis tasks. Method: We introduce SECQUE, the first dedicated benchmark for SEC filing analysis, comprising 565 expert-crafted questions spanning comparative analysis, ratio computation, risk assessment, and financial insight generation. We propose SECQUE-Judge, a multi-LLM collaborative judging framework that significantly improves alignment between automated and human evaluation (Cohen’s ΞΊ = 0.82), and systematically define and quantify LLMs’ capability boundaries in professional finance. Contribution/Results: Leveraging a structured evaluation protocol, expert annotations, and an integrated automatic assessment methodology, we comprehensively evaluate mainstream LLMs, exposing critical deficiencies in complex financial reasoning. SECQUE is fully open-sourced to advance reproducibility, comparability, and sustainable progress in financial AI evaluation.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Evaluation and Analysis

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
πŸ“ Abstract
We introduce SECQUE, a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis, ratio calculation, risk assessment, and financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human evaluations. Additionally, we provide an extensive analysis of various models' performance on our benchmark. By making SECQUE publicly available, we aim to facilitate further research and advancements in financial AI.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs in financial analysis tasks
Assessing performance on SEC filings questions
Advancing financial AI research with benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive benchmark for financial analysis tasks
Evaluation mechanism with multiple LLM-based judges
Publicly available to advance financial AI research
πŸ’Ό Related Jobs
No related jobs found.
N
Noga Ben Yoash
Microsoft Industry AI
M
Meni Brief
Microsoft Industry AI
O
Oded Ovadia
Microsoft Industry AI
G
Gil Shenderovitz
Microsoft Industry AI
M
Moshik Mishaeli
Microsoft Industry AI
R
Rachel Lemberg
Microsoft Industry AI
Eitam Sheetrit
Eitam Sheetrit
Ph.D., Software and Information Systems Engineering in Ben-Gurion University