Score
Designs and builds simulated users and virtual patients (including persona-driven help-seekers, auditor role-play agents, and embodied patient avatars) that produce longitudinal, multi-session interaction trajectories and coherent multi-turn utterances. These systems can be record-grounded (e.g., turning static records into dialogues), preserve factual and developmental consistency across sessions, and are engineered to support scalable scenario construction and large simulated populations.
This paper systematically investigates critical challenges in large language model (LLM)-driven role-playing (RP), focusing on character authenticity and personalization. It identifies three core problems: weak personality consistency, difficulty in behavior alignment with role specifications, and insufficient user engagement. To address these, the paper proposes the first four-dimensional technical taxonomy—spanning data curation, model alignment, agent architecture, and evaluation—highlighting dynamic personality modeling and higher-order consistency as pivotal research directions. Methodologically, it integrates prompt engineering, supervised and reinforcement fine-tuning, multi-agent collaboration, and hybrid human-automated multidimensional evaluation. Key contributions include: (1) the first structured, comprehensive research landscape map for RP; (2) an open-sourced, authoritative RP literature repository on GitHub; and (3) a reproducible benchmark evaluation framework. Collectively, these advances provide both theoretical foundations and practical paradigms for immersive AI-character interaction.
This work addresses the critical role of conversational user simulation in human-computer interaction, noting the absence of a systematic synthesis in existing literature. To bridge this gap, the paper proposes a unified classification framework grounded in large language models, integrating prior research along two key dimensions: user granularity and simulation objective. By structuring a comprehensive review around these axes, the study systematically examines core technical approaches, evaluation methodologies, and application scenarios. This structured analysis clarifies prevailing challenges and outlines promising future directions, thereby establishing a coherent research trajectory and theoretical foundation for advancing the field of conversational user simulation.
Existing patient simulators struggle to balance clinical authenticity with personality diversity, hindering the training and evaluation of large language models (LLMs) in multi-turn, context-aware doctor–patient dialogues. Method: We propose the first clinically grounded, four-dimensional persona modeling framework—encompassing personality traits, linguistic proficiency, medical history recall fidelity, and cognitive status—generating 37 composable, realistic patient personas from MIMIC-ED/IV real-world emergency department and intensive care data. Our approach integrates clinical knowledge graphs, multi-dimensional persona prompt engineering, and Llama 3.3–driven dialogue generation, validated by domain-expert physicians. Contribution/Results: Evaluated across eight state-of-the-art LLMs, our framework achieves high factual accuracy and persona consistency; blinded assessments by four clinicians confirm high clinical fidelity. The open-source, privacy-compliant system supports customizable medical education and standardized benchmarking—marking the first solution unifying clinical realism with scalable personality diversity.
Existing patient simulation methods often suffer from insufficient realism and limited controllability, frequently leading to excessive information disclosure and inadequate behavioral diversity. To address these limitations, this work proposes the PatientsWithPersonality (PWP) framework, which introduces the HEXACO six-dimensional personality model into virtual patient modeling for the first time. By explicitly parameterizing latent personality traits, PWP enables fine-grained control over dialogue style, cooperativeness, and information disclosure. Integrated with large language model generation and validated through automated scoring, PWP significantly enhances simulation realism. Clinical evaluations demonstrate that the generated dialogues closely resemble those produced by human actors, substantially outperforming current approaches while markedly reducing instances of information overdisclosure.
Current mental health conversational systems often rely on static prompts for patient simulators, resulting in homogeneous behaviors and incoherent symptom progression across multi-turn interactions. To address this limitation, this work proposes the DEPROFILE framework, which uniquely integrates real-world longitudinal clinical and life-event data into patient simulation. By fusing multi-source information into a unified patient profile and introducing a Chain-of-Change agent that transforms noisy event logs into structured temporal memories, DEPROFILE significantly enhances the realism, behavioral diversity, and temporal consistency of simulated patients. Extensive evaluations demonstrate that the proposed approach consistently outperforms state-of-the-art baselines across multiple large language model backbones.
This study addresses the lack of realism and diversity in existing Chinese clinical dialogue simulation data, which hinders effective evaluation of large language models (LLMs) in patient role-playing scenarios. To bridge this gap, the authors construct Ch-PatientSim, the first Chinese patient simulation dataset grounded in the Big Five personality dimensions, and propose a novel multi-stage patient role-playing framework that requires no model fine-tuning. By integrating few-shot generation, human verification, and staged dialogue mechanisms, the framework produces personalized and authentic doctor–patient conversations. Experimental results demonstrate that the proposed approach significantly outperforms existing methods across multiple patient simulation dimensions, effectively mitigating the common issues of overly formal responses and insufficient personality expression, thereby enhancing both the realism and linguistic diversity of simulated patient behaviors.
Current VR-based clinical communication training systems suffer from rigid, non-customizable content. To address this, we propose a customizable VR simulation framework integrating large language models (LLMs) with embodied conversational agents (ECAs). Grounded in user-centered research, we identified three core requirements—authentic clinical scenarios, intuitive interaction, and unpredictable dialogue—and designed the Virtual AI Patient Simulator (VAPS), enabling instructors to construct personalized training scenarios without coding. Our framework overcomes traditional VR content-generation bottlenecks, significantly enhancing real-world scenario adaptability and pedagogical scalability beyond controlled lab settings. Empirical evaluation confirms that VAPS delivers highly immersive, realistic clinician–patient dialogues, effectively supporting diverse curricular objectives. (138 words)
This work addresses the significant challenge of simulating clinical patient trajectories, which are shaped by complex biological and social factors, thereby hindering advances in personalized medicine and virtual clinical trials. To this end, we leverage over 200 million real-world electronic health records to develop the first large-scale, pre-trained generative simulator capable of modeling the probabilistic distribution of future clinical events, laboratory results, and their temporal dynamics based solely on a patient’s historical data. The generated trajectories exhibit high fidelity to real-world observations, with incidence rates, lab values, and temporal patterns closely matching empirical data. Notably, the observed-to-expected ratios for diverse clinical outcomes consistently approximate 1.0, demonstrating the model’s effectiveness and potential for high-fidelity patient trajectory simulation.
This study addresses a critical limitation in existing large language model–driven client simulators for psychotherapy, which often produce overly compliant responses lacking the resistance and causal depth characteristic of real clinical interactions. To enhance ecological validity, the authors propose a novel client simulation framework grounded in clinical theory, integrating the 5Ps case formulation approach with a dynamic trust mechanism. This framework employs a dynamic memory layer to track the therapeutic alliance and introduces a trust threshold to modulate emotion–behavior modeling, thereby generating clinically coherent responses. Evaluated across 40 diverse clinical scenarios, the method produces highly plausible and varied client profiles, significantly outperforming baseline models in both the diversity of resistance expression and behavioral authenticity, thus advancing the clinical fidelity of simulated clients.
This study addresses the evaluation bottleneck caused by the high cost and limited scalability of real-user studies by proposing a role-based universal user simulation framework. By integrating role datasets with application interfaces to capture interaction trajectories, this method establishes a pluggable and parallelizable automated evaluation workflow. Empirical validation across diverse scenarios, including questionnaires, chatbots, and web applications, demonstrates that the framework supports plug-and-play testing and effectively overcomes traditional scalability limitations. Consequently, it generates reproducible, high-quality user feedback, offering an efficient and scalable paradigm for evaluating interactive applications.
本文针对面试对话系统测试所需大量人力的问题,提出了一种利用大型语言模型自动生成具有多样化沟通风格的用户模拟器人格的方法。
This work addresses the lack of standardization in existing patient simulation methods, which exhibit incompatible data formats, prompt templates, and evaluation metrics, thereby severely hindering reproducibility and fair comparison. To overcome this limitation, we propose the first standardized and modular patient simulation framework that unifies the definition, composition, and deployment of simulated patients, enabling cross-method evaluation and seamless integration of custom metrics. Built upon a large language model–based role-playing architecture, the framework offers high flexibility and interoperability with diverse simulation strategies and assessment mechanisms. We demonstrate its effectiveness by successfully reproducing multiple representative approaches and rapidly developing two novel simulators, thereby validating the framework’s extensibility and capacity to accelerate research and development. The code is publicly released.