MemTrial: Learning When to Trust Memory in LLM Portfolio Agents

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem that large language model (LLM) investment agents underperform equal-weight benchmarks due to memory attribution errors induced by market noise. To overcome this limitation, we propose a "Memory Trial" mechanism. Specifically, this approach compares decision drafts generated with and without experiential knowledge, leveraging Banzhaf values to achieve precise experience attribution. Furthermore, it incorporates fractional factor design and hierarchical Bayesian inference to facilitate robust experience trust learning. Extensive evaluations across multiple benchmarks demonstrate that the proposed method significantly outperforms existing baselines, achieving a 21.2% improvement in utility while maintaining minimal downside risk. Overall, this work provides a reliable paradigm for experience utilization in LLM-driven financial decision-making.
📝 Abstract
Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/$N$) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/$N$. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2\% below 1/$N$ on PortBench and InvestorBench, against 15--38\% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2\%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.
Problem

Research questions and friction points this paper is trying to address.

Portfolio Management
LLM Agents
Credit Assignment
Experience Learning
Market Noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM portfolio agents
Banzhaf value
fractional factorial design
hierarchical Bayesian model
credit assignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guanghao Wu
University of Technology Sydney
Z
Zhuo Cai
University of Technology Sydney
Shoujin Wang
Shoujin Wang
University of Technology Sydney
Data ScienceMachine LearningRecommender SystemMisinformationData Science Application