Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the susceptibility of large language models (LLMs) to numerical reasoning errors and insufficient faithfulness when generating explanations for stock predictions. To overcome these limitations, this work proposes a progressive reasoning externalization framework. The method integrates XGBoost with temporal SHAP to extract numerical features and employs historical analogies to transform complex numerical relationships into deterministic descriptions, which subsequently guide the Qwen3 model in producing verifiable financial narratives. Experimental results demonstrate that the proposed framework elevates evidence faithfulness to 0.996 while significantly improving the accuracy of temporal and logical relations as well as human-rated usefulness scores. These findings indicate that the framework offers an effective solution for enhancing both the interpretability and reliability of LLM-driven financial decision-making.
πŸ“ Abstract
In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Stock Return Prediction
Explainable AI
Narrative Faithfulness
Reasoning Externalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reasoning Externalization
Temporal SHAP
Evidence Faithfulness
Historical Regime Analogs
LLM Narrative Framework
πŸ”Ž Similar Papers
No similar papers found.
S
Sujung Kim
Department of Industrial Data Engineering, Hanyang University, Republic of Korea
S
Seung Hwan Cho
Department of Industrial Data Engineering, Hanyang University, Republic of Korea
S
Sangjin Park
School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea
Young-Min Kim
Young-Min Kim
Associate Professor, Hanyang University
Machine LearningInformation ExtractionProbabilistic ModelsNatural Language Processing