🤖 AI Summary
This study addresses the distortion of evaluation metrics and the absence of methodological benchmarks in long-document financial narrative summarization, utilizing annual reports from the London Stock Exchange as its empirical foundation. Methodologically, it systematically compares various pretrained Transformer models with extractive techniques and introduces BRUGEscore, a novel metric integrating ROUGE-2 and BERTScore. Statistical significance testing and adversarial analysis are further employed to enhance experimental reliability. The results demonstrate that existing metrics exhibit substantial limitations, whereas BRUGEscore more faithfully reflects model summarization capabilities. By exposing the deficiencies of conventional evaluation metrics, this research establishes a robust evaluation benchmark for long-document financial summarization tasks.
📝 Abstract
There are more than 2,000 listed companies on the UK's London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a variety of summarisation methods on a set of different pre-trained transformers with different extraction techniques. In addition, we considered multiple evaluation metrics in order to investigate their differing behaviour and applicability on a dataset from the Financial Narrative Summarisation (FNS 2020) shared task, which is composed of annual reports published by firms listed on the London Stock Exchange and their corresponding summaries. We hypothesise that some evaluation metrics do not reflect true summarisation ability and propose a novel BRUGEscore metric, as the harmonic mean of ROUGE-2 and BERTscore. Finally, we perform a statistical significance test on our results to verify whether they are statistically robust, alongside an adversarial analysis task with three different corruption methods.