🤖 AI Summary
This study addresses the limitations of offline-trained large language model (LLM) psychological counseling systems, which struggle to accommodate individual differences and lack online learning capabilities. To this end, we propose a test-time self-evolving counseling agent framework. Methodologically, a hierarchical Bayesian skill policy is designed to enable personalized intervention selection, integrated with inter-session listwise preference optimization to enhance response quality. Furthermore, state-conditioned ordinal credit assignment is employed to provide feedback signals. These three components synergistically facilitate cross-client preference learning and individualized skill posterior updates. Evaluated on the PsychEval benchmark, the proposed framework achieves an overall score of 7.684, outperforming all ablated variants and thereby validating both the effectiveness and the synergistic contributions of each component.
📝 Abstract
Mental health disorders affect a substantial proportion of the global population, yet a persistent shortage of trained practitioners leaves the majority without adequate care. Large language model (LLM)-based counselors present a promising direction for delivering scalable conversational psychological support. Offline model training alone leaves limited room to adapt to individual clients or to learn from ongoing therapeutic interaction at test time. We introduce PsyEvo, an LLM-based counseling framework that enables both client-specific personalization and response-policy improvement at test time through three components: Hierarchical Bayesian Skill Policy (HBSP) personalizes what intervention to apply by maintaining a per-client skill posterior updated from session feedback; Inter-session Listwise Preference Optimization (LiPO) improves how the selected skill is expressed by updating a shared response adapter from cross-client preference evidence; and State-conditioned Ordinal Credit Assignment (SOCA) supplies candidate preferences and trajectory credit to the two components through consistency-checked comparisons and ordinal projection. In simulated-client evaluation with shared online cohort adaptation, PsyEvo obtains 7.684 Overall on PsychEval and exceeds every component variant in each of three matched runs. Removing individual components lowers mean overall score by 0.138--0.171 under the shared configuration, supporting conditional contributions within the complete scaffold. Our code is available at https://github.com/Lingxi-mental-health/PsyEvo