FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

πŸ“… 2026-08-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the gap between the fluency of financial reports generated by current large language models and the stringent institutional requirements for identity consistency, content completeness, data discipline, and visualization. To bridge this gap, the authors introduce FinReportBench, the first multimodal evaluation benchmark tailored to institutional standards, encompassing 35 fine-grained criteria, along with a reproducible expert-based pairwise preference assessment framework. By leveraging 244 bilingual tasks for skill distillation and integrating multimodal evidence fusion, decision-boundary auditing, and hierarchical input design, the approach transforms failure modes into reusable generation and self-audit constraints. Experiments demonstrate that this method improves average G1 scores by 33.85 points and G2 scores by 13.83 points across five mainstream models, significantly enhancing identity consistency and institutional completeness while preserving baseline deliverability.
πŸ“ Abstract
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert-grounded benchmark for measuring and improving institution-grade financial report generation. Expert review reveals recurring gaps in report identity, institutional components, source discipline, and visual delivery. We derive a 35-item rubric through expert partial orders, multimodal evidence, and audits of decision boundaries, covering deliverability, report identity, and institutional completeness. Starting from 10,000 balanced Chinese and English financial-research source records, we curate 244 bilingual tasks across three research objects and two input tiers. Each task separates the public query, reconstructed research trajectory, and hidden source packet. Three independent judge families reproduce the expert partial order at near-ceiling rates, showing that bounded, observable criteria support reliable evaluation. Across nine model families, basic deliverability is nearly saturated, while report identity and institutional completeness remain the primary bottlenecks. The largest cross-model gaps concern generation-trace control, information density, and data discipline rather than basic report framing. We then use benchmark-guided skill distillation to turn recurrent failures into reusable generation and self-review constraints. Across five model families, the evolved skill improves mean G1 by 33.85 points and mean G2 by 13.83 points over paired no-skill runs while preserving G0 for every pair. Code and benchmark artifacts are available at https://github.com/MisterBrookT/finreportbench.
Problem

Research questions and friction points this paper is trying to address.

financial report generation
institution-grade evaluation
report identity
institutional completeness
deliverability
Innovation

Methods, ideas, or system contributions that make the work stand out.

FinReportBench
institution-grade report generation
expert-grounded benchmark
skill distillation
multimodal evaluation