π€ AI Summary
Large language models (LLMs) lack reliable, domain-specific evaluation benchmarks for real-world financial analysis tasks. Method: We introduce SECQUE, the first dedicated benchmark for SEC filing analysis, comprising 565 expert-crafted questions spanning comparative analysis, ratio computation, risk assessment, and financial insight generation. We propose SECQUE-Judge, a multi-LLM collaborative judging framework that significantly improves alignment between automated and human evaluation (Cohenβs ΞΊ = 0.82), and systematically define and quantify LLMsβ capability boundaries in professional finance. Contribution/Results: Leveraging a structured evaluation protocol, expert annotations, and an integrated automatic assessment methodology, we comprehensively evaluate mainstream LLMs, exposing critical deficiencies in complex financial reasoning. SECQUE is fully open-sourced to advance reproducibility, comparability, and sustainable progress in financial AI evaluation.
π Abstract
We introduce SECQUE, a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis, ratio calculation, risk assessment, and financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human evaluations. Additionally, we provide an extensive analysis of various models' performance on our benchmark. By making SECQUE publicly available, we aim to facilitate further research and advancements in financial AI.