A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on costly expert annotations, challenges in long-context processing, and prevalent hallucinations lacking trustworthy benchmarks when large language models (LLMs) analyze financial reports. We propose an expert-annotation-free numerical evidence evaluation method. By constructing an automated data pipeline to generate an S&P 500 benchmark dataset, we introduce the first unsupervised groundedness metric and a "conscious incompetence" detection mechanism to identify scenarios with insufficient evidence. Experiments reveal that while LLMs exhibit strong grounding capabilities, their factual correctness remains limited, rendering them prone to hallucination when critical information is missing. This work establishes a novel paradigm for quantifying and mitigating LLM hallucinations in the financial domain.
📝 Abstract
Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the S&P 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounded Analysis
Numeric Evidence Evaluation
Automated Dataset Construction
Conscious Incompetence
Earnings Call Transcripts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yingzhu Zhao
Global Decision Science, American Express
V
Vlad Pandelea
Global Decision Science, American Express
H
Han Yuan
Global Decision Science, American Express
B
Bo Hu
Global Decision Science, American Express
W
Wuqiong Luo
Global Decision Science, American Express
L
Li Zhang
Global Decision Science, American Express
Z
Zheng Ma
Singapore Decision Science Center of Excellence, American Express