How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment

📅 2026-07-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional financial sentiment analysis, which predominantly relies on news texts and uses stock returns as the sole predictive label, thereby overlooking the predictive power of regulatory filings—particularly the Item 1A risk factors section of 10-K reports—for market volatility. The authors propose a supervised dictionary learning approach to construct sentiment indicators across three aggregation levels—firm, portfolio, and industry—using 1,383 10-K filings from 94 Nasdaq-100 technology companies, with both returns and volatility as target variables. Results reveal that full-text sentiment performs better at higher aggregation levels (industry and portfolio), whereas Item 1A excels at the individual firm level. In contrast, the conventional Loughran-McDonald dictionary exhibits significant negative correlations. These findings underscore the necessity of supervised learning for regulatory text and highlight the interaction between textual scope and aggregation level in financial sentiment modeling.
📝 Abstract
Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored. We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm. Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006--2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realised market outcomes, and qualitative lexical content. Full-filing text produces more accurate sentiment at the sector and portfolio level for both targets, but this reverses at the individual-firm level, where the narrower Item 1A section performs better -- an effect we attribute to the interaction between document volume and the amount of independent training signal available at each level of aggregation. A Loughran-McDonald dictionary baseline is consistently, strongly negatively correlated with price at every level tested, underscoring the value of a supervised approach for regulatory disclosure text. These findings, and the design choices they motivate, establish the sentiment-generation methodology underlying a subsequent, larger-scale, multi-source system.
Problem

Research questions and friction points this paper is trying to address.

10-K filings
financial sentiment
risk factors
volatility
text aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

supervised lexicon learning
10-K filings
risk-factor sentiment
volatility prediction
aggregation-dependent performance
🔎 Similar Papers
2023-06-17Social Science Research NetworkCitations: 25