PreFER: Interactive Robo-Advisor with Scoring Mechanism

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of traditional robo-advisors in accurately learning personalized risk preferences from user ratings and dynamically adapting to market conditions. To this end, we propose an interactive intelligent investment advisory framework based on inverse reinforcement learning. Methodologically, we introduce a discrete-time predictable forward exploration reward (PreFER) process, which integrates von Neumann acceptance-rejection sampling with constant absolute risk aversion (CARA) utility theory to efficiently transform noisy ratings into optimal policy distributions. Experimental results demonstrate that the proposed framework can accurately identify clients’ risk aversion levels despite noisy feedback. Furthermore, it consistently generates dynamic investment recommendations aligned with the learned preferences as market conditions evolve.
πŸ“ Abstract
We propose an interactive robo-advising framework that learns personalized risk preferences from scores provided by clients. The resulting preference-learning problem is closely related to inverse reinforcement learning (IRL), as the robo-advisor infers the client's latent reward specification from feedback. The robo-advisor interacts with clients iteratively as follows. At each interaction time, the advisor generates investment advice based on the optimal policy distribution derived from an inferred personalized risk preference. The client scores the advice. The advisor updates its assessment of the client's risk preference based on the feedback. This learning procedure motivates us to investigate discrete-time Predictable Forward Exploratory Reward (PreFER) processes and derive an exploratory investment strategy. By interpreting the score as the acceptance probability of a piece of advice, our inverse learning procedure learns the client's exploratory investment distribution using the acceptance-rejection method pioneered by von Neumann. Under CARA preferences, we show that, even though the scores contain noise, the robo-advisor can identify the client's current risk aversion after a sufficiently large number of interactions. The PreFER process then carries the learned preference forward and generates future recommendations under updated market conditions.
Problem

Research questions and friction points this paper is trying to address.

Robo-Advisor
Risk Preference Learning
Inverse Reinforcement Learning
Personalized Investment
Scoring Mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Robo-Advisor
Inverse Reinforcement Learning
Predictable Forward Exploratory Reward (PreFER)
Acceptance-Rejection Method
Risk Preference Learning
πŸ’Ό Related Jobs
No related jobs found.
Y
Yuwei Wang
School of Mathematics, Shanghai University of Finance and Economics, Shanghai 200433, China
Hoi Ying Wong
Hoi Ying Wong
Chinese University of Hong Kong
FinanceOption pricingportfolio theorystochastic processstochastic control