MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement

📅 2026-03-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language models (LLMs) in financial forecasting, where reliance on memorized training data often generates spurious alpha signals that lead to out-of-sample performance collapse. To mitigate this, the authors propose MemGuard-Alpha, a zero-cost, signal-level framework for detecting and filtering memory-contaminated predictions. The approach integrates five membership inference attacks into a composite memory contamination score (MCS) and introduces a cross-model disagreement detection algorithm (CMMD) that leverages differences in model training cutoff dates to identify compromised signals in real time. Empirical evaluation across seven LLMs and 50 S&P 100 constituents demonstrates that filtered signals achieve a 49% higher Sharpe ratio (reaching 4.11) and a sevenfold increase in average daily return (14.48 bps versus 2.13 bps), substantially enhancing the robustness of quantitative trading strategies.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Large Vision Models

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) are increasingly used to generate financial alpha signals, yet growing evidence shows that LLMs memorize historical financial data from their training corpora, producing spurious predictive accuracy that collapses out-of-sample. This memorization-induced look-ahead bias threatens the validity of LLM-based quantitative strategies. Prior remedies -- model retraining and input anonymization -- are either prohibitively expensive or introduce significant information loss. No existing method offers practical, zero-cost signal-level filtering for real-time trading. We introduce MemGuard-Alpha, a post-generation framework comprising two algorithms: (i) the MemGuard Composite Score (MCS), which combines five membership inference attack (MIA) methods with temporal proximity features via logistic regression, achieving Cohen's d = 18.57 for contamination separation (d = 0.39-1.37 using MIA features alone); and (ii) Cross-Model Memorization Disagreement (CMMD), which exploits variation in training cutoff dates across LLMs to separate memorized signals from genuine reasoning. Evaluated across seven LLMs (124M-7B parameters), 50 S&P 100 stocks, 42,800 prompts, and five MIA methods over 5.5 years (2019-2024), CMMD achieves a Sharpe ratio of 4.11 versus 2.76 for unfiltered signals (49% improvement). Clean signals produce 14.48 bps average daily return versus 2.13 bps for tainted signals (7x difference). A striking crossover pattern emerges: in-sample accuracy rises with contamination (40.8% to 52.5%) while out-of-sample accuracy falls (47% to 42%), providing direct evidence that memorization inflates apparent accuracy at the cost of generalization.
Problem

Research questions and friction points this paper is trying to address.

memorization
financial forecasting
look-ahead bias
large language models
alpha signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

membership inference attack
memorization filtering
cross-model disagreement
financial alpha signals
LLM-based forecasting
A
Anisha Roy
Department of Electronics and Communication Engineering, Jaypee Institute of Information Technology, Noida, India
Dip Roy
Dip Roy
Indian Institute of Technology, Patna
Explainable AIMechanistic Interpretibility