evidence-grounded response generation

Designs, builds, or evaluates conversational response systems that retrieve and integrate evidence-based therapeutic literature and the current conversational context to produce clinically aligned, context-aware therapeutic replies; implements context-aware prompting and safety mechanisms to enforce clinical constraints and reduce harmful or non-evidence-based guidance.

evidence-groundedresponsegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Streamlining Biomedical Research with Specialized LLMs

Apr 15, 2025
LC
Linqing Chen
🏛️ PatSnap Co., LTD.

To address low knowledge acquisition efficiency and unverifiable responses in biomedical and pharmaceutical R&D, this paper proposes an interactive system featuring deep synergy between domain-specific large language models (LLMs) and retrieval mechanisms. Methodologically, it integrates a fine-tuned biomedical LLM, hybrid retrieval (semantic + keyword-based), multimodal response orchestration, and a cross-source corroboration reasoning framework—enabling context-aware generation and dynamic fusion of text, images, and tabular data. Its key innovation is the first-of-its-kind bidirectional retrieval-generation verification architecture, ensuring response traceability, multi-source evidence alignment, and high-fidelity real-time dialogue. Experimental results demonstrate significant improvements in question-answering accuracy and decision-making efficiency during R&D phases. The system has been deployed as a structured knowledge service platform for pharmaceutical enterprises and biomedical researchers.

Enhancing biomedical research efficiency with specialized LLMsFacilitating real-time, high-fidelity human-computer interactionsImproving response precision via advanced information retrieval

Current large language models lack structured evaluation of core therapeutic principles in mental health conversations, compromising clinical appropriateness. To address this gap, this work introduces FAITH-M—the first expert-annotated benchmark grounded in established therapeutic principles—and CARE, a multi-stage evaluation framework that enables precise assessment of AI therapist responses through fine-grained ordinal scoring, context-aware analysis, contrastive example retrieval, and chain-of-thought knowledge distillation. Using Qwen3 as the backbone model, CARE achieves an F1 score of 63.34, representing a 64.26% improvement over baseline methods, and demonstrates strong robustness across diverse datasets and expert evaluations.

AI evaluationclinical appropriatenessmental health conversation

This study addresses the limited contextual understanding and affective awareness of large language models (e.g., GPT-2) in mental health support dialogues. We propose a novel method integrating structured input reconstruction with multi-component reinforcement learning (MRL), explicitly modeling user utterances, dialogue history, and fine-grained emotional states. A multi-objective reward function jointly optimizes contextual coherence, affective consistency, and clinical appropriateness, enabling synergistic supervised fine-tuning and RL-based training. Experiments demonstrate a significant improvement in emotion recognition accuracy—from 66.96% to 99.34%—and consistent gains across generation metrics (BLEU, ROUGE, METEOR) over all baselines. Human evaluation by LLM-based annotators confirms high response relevance and clinical plausibility. To our knowledge, this is the first work to systematically incorporate MRL into psychotherapeutic dialogue generation, substantially enhancing the model’s joint situational–affective modeling capability.

Developing AI systems that align with professional therapist standards and emotional accuracyEnhancing emotional awareness in therapeutic dialogue generation for mental health supportImproving contextual understanding of pre-trained language models for therapy responses

MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation

Aug 26, 2025
EL
Ernest Lim
🏛️ Ufonia Limited | University of York | NHS Improvement Academy

Current clinical dialogue systems lack systematic evaluation methodologies tailored for safety-critical scenarios. To address this, we propose MATRIX—a novel framework that tightly integrates structured safety engineering with scalable dialogue AI assessment, establishing a comprehensive safety evaluation system encompassing clinical risk scenarios, behavioral norms, and failure modes. Our key contributions are threefold: (1) BehvJudge, an expert-validated evaluator achieving F1 = 0.96 in hazardous utterance detection across 240 clinical dialogues—surpassing human clinicians; (2) PatBot, a high-fidelity patient simulation agent covering 14 risk categories across 10 clinical domains, validated for realism in 2,100 simulated interactions; and (3) a regulatory-aligned safety auditing workflow. MATRIX establishes a reproducible, scalable, and interpretable automated assessment paradigm for safety verification of medical AI systems.

Detecting safety-relevant dialogue failures systematicallyEvaluating safety risks in clinical dialogue systemsSimulating realistic patient interactions for safety testing

Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for Psychotherapy

Nov 11, 2024
XS
Xin Sun
🏛️ University of Amsterdam | Centrum Wiskunde & Informatica (CWI) | Tilburg University | Southeast University | Vrije Universiteit Amsterdam | Utrecht University

Large language models (LLMs) in psychotherapy chatbots suffer from weak controllability, low therapeutic adherence, and excessive reliance on handcrafted scripts. Method: We propose Script-Strategy Aligned Generation (SSAG), a novel paradigm that aligns LLM outputs with clinical strategies—e.g., cognitive behavioral therapy (CBT) frameworks—during dynamic generation, rather than merely matching static scripts. SSAG integrates instruction-tuned prompting, supervised fine-tuning, and strategy-guided multi-stage alignment to yield an interpretable and controllable generation mechanism. Contribution/Results: In a 10-day field study, SSAG achieved clinically viable performance in therapeutic adherence, user trust, and intervention efficacy—matching fully scripted systems and significantly outperforming rule-based baselines. This work establishes a new pathway for digital mental health interventions that balances clinical rigor, safety, and development efficiency.

Aligning LLMs with expert-crafted psychotherapy scripts for controllabilityEnhancing LLM adherence to therapeutic principles while maintaining flexibilityReducing reliance on rigid rule-based systems in therapeutic chatbots

Latest Papers

What's happening recently
View more

This study addresses the challenge clinicians face in applying lengthy evidence-based guidelines during fast-paced primary care consultations. To bridge this gap, the authors propose a multi-stage reasoning prompting strategy leveraging the large language model Gemini 2.5 to automatically generate clinically relevant questions—rather than direct answers—that align closely with established guidelines from real patient–physician dialogues. The approach, implemented via zero-shot and multi-stage prompting, was evaluated on 80 authentic consultation transcripts and rigorously assessed by six senior physicians over more than 90 hours. Results demonstrate that the generated questions exhibit strong clinical relevance and guideline adherence, significantly reducing physicians’ cognitive load and enhancing the practical implementation of evidence-based medicine in frontline clinical practice.

Clinical decision supportEvidence-based medicinePhysician-patient dialogue

This work addresses the scarcity of mental health counselors and delayed response times in online psychological support, particularly in low-resource language settings such as Hebrew and Arabic, where intelligent tools capable of delivering expert-level crisis intervention are lacking. The authors propose CARE, a novel framework that leverages expert-validated, high-quality crisis intervention dialogues to perform domain-aligned fine-tuning of open-source large language models, enabling context-aware response generation grounded in full conversation history. Experimental results demonstrate that CARE significantly outperforms general-purpose large language models in both semantic fidelity and adherence to evidence-based intervention strategies, producing responses closely aligned with those of professional counselors. This approach effectively enhances the quality and efficiency of psychological support in linguistically underserved populations.

counselor alignmentcrisis interventionlarge language models

This work addresses the challenge of efficiently integrating heterogeneous, multi-source data in biomedical research by proposing and implementing an intelligent scientific assistant powered by large language models. The system unifies multimodal data—including scientific literature, knowledge graphs, chemical databases, and clinical trial records—through semantic retrieval, enabling both question-answering and multi-step reasoning interactions. It incorporates an evidence-tracing mechanism to ensure interpretability and auditability of its outputs. As the first system to achieve cross-source semantic integration and traceable reasoning in pharmaceutical R&D, it has been deployed across AstraZeneca’s global research infrastructure, significantly enhancing researchers’ information retrieval efficiency and their capacity for automated exploration of complex drug discovery tasks.

biomedical researchdata integrationevidence retrieval

This study addresses the global shortage of evidence-based psychotherapy resources—even in high-income regions, where long wait times persist—by proposing a large language model–based embodied conversational agent. The system uniquely integrates multilevel psychological analysis, process-oriented therapeutic principles, and embodied interaction, leveraging retrieval-augmented generation, emotion recognition, psychological flexibility assessment, and synchronized speech animation to deliver real-time, clinically safe, and evidence-informed responses. Evaluated under a GPT-5.2 configuration, the agent outperformed human therapist responses in comprehension, interpersonal effectiveness, collaboration, and therapeutic adherence, and received endorsement from eleven licensed psychotherapists. This work establishes a novel, supervisable, and safety-controlled paradigm for AI-delivered psychological support.

evidence-based therapymental health supportpsychotherapy access

This study addresses the lack of fine-grained, interpretable evaluation mechanisms for assessing how counselors respond to client resistance—a critical gap that hinders skill development in psychotherapy. To bridge this gap, the authors propose a novel, theory-driven multidimensional evaluation framework that decomposes counselor responses in resistance scenarios into four distinct communication mechanisms. Leveraging expert-annotated data, they perform full-parameter instruction tuning on Llama-3.1-8B-Instruct to jointly model response quality scoring and explanatory rationale generation. The resulting model achieves 77–81% F1 in identifying the quality of different communication mechanisms, significantly outperforming GPT-4o and Claude-3.5-Sonnet. Moreover, its generated explanations receive near-perfect expert ratings (2.8–2.9/3.0) and are empirically validated to effectively enhance the resistance-handling competence of 43 practicing counselors.

client resistancecounselor feedbackexplainable AI

Hot Scholars

LL

Lewei Lu

Research Director (We're Hiring, luotto@sensetime.com) @ SenseTime Research
Computer VisionDeep Learning
DG

Diandian Guo

The Chinese University of Hong Kong
Deep learning
SH

Simon Hoermann

Associate Professor at University of Canterbury
HCIHealth TechnologiesGames User ResearchPositive Computing
HC

Hyungjoo Chae

Georgia Institute of Technology
GUI AgentDigital AgentLLM Agent
JZ

Jingyu Zhang

WNLO Huazhong University of Science and Technology
optical