🤖 AI Summary
Large language models (LLMs) frequently exhibit intent drift, contextual incoherence, and factual hallucinations in dialogue, undermining their reliability in real-world applications. To address this, we conduct a rapid systematic review guided by the PRISMA framework and the PICO strategy. This work introduces— for the first time—a taxonomy of dialogue alignment techniques structured along the LLM lifecycle: inference-time, post-training, and reinforcement learning stages. We particularly highlight inference-time interventions—including prompt engineering, self-verification, and retrieval-augmented generation—which improve intent consistency, contextual groundedness, and hallucination suppression *without* model retraining. Empirical findings demonstrate that these methods offer high efficiency, practical deployability, and flexibility across diverse deployment scenarios. Our taxonomy and analysis thus provide both a theoretically grounded framework and an actionable technical pathway for enhancing dialogue reliability in production LLM systems.
📝 Abstract
Large language models (LLMs) may generate outputs that are misaligned with user intent, lack contextual grounding, or exhibit hallucinations during conversation, which compromises the reliability of LLM-based applications. This review aimed to identify and analyze techniques that align LLM responses with conversational goals, ensure grounding, and reduce hallucination and topic drift. We conducted a Rapid Review guided by the PRISMA framework and the PICO strategy to structure the search, filtering, and selection processes. The alignment strategies identified were categorized according to the LLM lifecycle phase in which they operate: inference-time, post-training, and reinforcement learning-based methods. Among these, inference-time approaches emerged as particularly efficient, aligning outputs without retraining while supporting user intent, contextual grounding, and hallucination mitigation. The reviewed techniques provided structured mechanisms for improving the quality and reliability of LLM responses across key alignment objectives.