🤖 AI Summary
This study addresses the reliance on costly expert annotations, challenges in long-context processing, and prevalent hallucinations lacking trustworthy benchmarks when large language models (LLMs) analyze financial reports. We propose an expert-annotation-free numerical evidence evaluation method. By constructing an automated data pipeline to generate an S&P 500 benchmark dataset, we introduce the first unsupervised groundedness metric and a "conscious incompetence" detection mechanism to identify scenarios with insufficient evidence. Experiments reveal that while LLMs exhibit strong grounding capabilities, their factual correctness remains limited, rendering them prone to hallucination when critical information is missing. This work establishes a novel paradigm for quantifying and mitigating LLM hallucinations in the financial domain.
📝 Abstract
Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the S&P 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.