Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that the unstructured nature of SEC 10-K filings leads to over 70% missing corporate financial data, introducing biases in quantitative analysis and systematic underrepresentation of smaller firms. To overcome this, we propose a generative AI pipeline leveraging large language models (Llama-3 and Qwen-2.5), introducing a strategy that aligns model parameter scale with document complexity. Combined with zero-shot prompt engineering, this approach enables precise extraction of multidimensional financial attributes across sections and footnotes. Our method effectively mitigates the fragility of traditional information extraction techniques, achieving strong performance in a zero-shot setting—notably an F1 score of 83.33% for tabular data using Qwen—while substantially eliminating small-firm data bias. This work establishes a reliable paradigm for quantitative research on unstructured financial texts.
📝 Abstract
SEC 10-K filings contain substantial financial information that is not consistently captured in structured datasets, creating a missing-data problem affecting over 70% of firms and half of total market capitalization. This can disproportionately bias quantitative analysis against smaller firms, which may be excluded due to limited available data. Traditional financial extraction methods such as Regular Expressions (Regex) and BERT, have been widely used. However, they are highly brittle when parsing complex SEC 10-K filings, which leads to data that is existent in the files being lost since these methods do not consider that a data attribute could be located in a different section or a footnote. This study evaluates several Large Language Models (LLMs), including Llama-3 8B, Qwen-2.5 14B, and Llama-3.3 70B, to figure out individual model strengths and weaknesses when extracting specific attributes from SEC 10-K text. The extraction quality was evaluated across four financial variables of varying structural complexity: Cash and Cash Equivalents (tabular), Short-Term Debt (hybrid), Credit Facilities (narrative), and Research and Development (hybrid). Results show that while smaller models like Llama-3 8B experience performance degradation under complex negative prompting, aligning parameter scale with document complexity yields high zero-shot accuracy. Qwen-2.5 14B excels as a tabular specialist with an 83.33% F1 score on Cash, whereas Llama-3.3 70B effectively navigates dense narrative footnotes, achieving a 76.92% F1 score on R&D. This scalable framework addresses critical information gaps in quantitative finance datasets and eliminates missing-data bias through a more thorough analysis of the SEC 10-K files.
Problem

Research questions and friction points this paper is trying to address.

Financial Data Extraction
Missing Data
SEC 10-K Filings
Quantitative Finance
Data Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative AI Pipeline
Large Language Models
Zero-shot Extraction
SEC 10-K Filings
Missing Data Bias
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.