Multi-Level Testing of Conversational AI Systems

πŸ“… 2026-02-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge that existing testing methodologies struggle to effectively validate the behavior of conversational AI systems under interactive and multi-granularity integration scenarios. To this end, it proposes the first hierarchical testing framework specifically designed for conversational AI, systematically covering verification across multiple levelsβ€”from language and AI components, through single-agent systems, to multi-agent configurations. The approach integrates principles from software testing theory, architectural analysis of dialogue systems, and AI behavior validation techniques to construct a layered testing strategy. Experimental results demonstrate that the framework significantly enhances system reliability and consistency across different integration layers, offering a scalable and structured testing paradigm for complex conversational AI systems.

Technology Category

Natural Language Processing: Conversational AI/Dialog SystemsMultiagent Systems: Agent CommunicationCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Search and Retrieval-Augmented AI: Assisted, interactive, and conversational searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
πŸ“ Abstract
Conversational AI systems combine AI-based solutions with the flexibility of conversational interfaces. However, most existing testing solutions do not straightforwardly adapt to the characteristics of conversational interaction or to the behavior of AI components. To address this limitation, this Ph.D. thesis investigates a new family of testing approaches for conversational AI systems, focusing on the validation of their constituent elements at different levels of granularity, from the integration between the language and the AI components, to individual conversational agents, up to multi-agent implementations of conversational AI systems
Problem

Research questions and friction points this paper is trying to address.

Conversational AI
Testing
Multi-level validation
AI components
Conversational agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-level testing
conversational AI systems
AI component validation
multi-agent systems
conversational interfaces
E
Elena Masserini
University of Milano-Bicocca, Milan, Italy