ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

๐Ÿ“… 2026-08-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing evaluation methods struggle to capture the mechanisms through which conversational investment advisors influence usersโ€™ long-term investment behaviors. This work proposes ShiJianBench, the first multi-agent investor simulator that integrates motivation-driven decision-making, state evolution, and dialogue updates, establishing an offline evaluation framework grounded in historical market feedback. The framework enables fine-grained, trajectory-level assessment via behavioral calibration, long-term memory modeling, and compliance-aware gating strategies. Experiments on Chinese mutual fund market data from 2021 to 2026 demonstrate that the identified top-performing LLM-based advisors significantly outperform baseline systems in both personalized content generation and long-term investment outcomes.
๐Ÿ“ Abstract
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.
Problem

Research questions and friction points this paper is trying to address.

conversational investment advisors
long-horizon evaluation
investor behavior
trajectory-aware evaluation
advisor language
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon evaluation
multi-agent investor simulator
trajectory-aware assessment
conversational investment advisors
dialogue-grounded updates