🤖 AI Summary
This work addresses the lack of domain-specific, real-world evaluation benchmarks for automatic speech recognition in the financial domain by introducing a large-scale dataset comprising 498 hours of full earnings call recordings and 46 hours of industry-balanced excerpts. For the first time, this benchmark provides fine-grained metadata—including speaker roles, meeting structure, and industry labels—and employs a standardized alignment pipeline to produce high-quality transcripts. Building upon this resource, the authors establish reproducible baseline systems using Whisper and Parakeet-TDT, enabling multidimensional, industry-aware, and role-sensitive evaluation beyond conventional word error rate metrics. This dataset fills a critical gap in evaluation resources for financial-domain speech recognition research.
📝 Abstract
We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.