π€ AI Summary
This work addresses the challenge that existing testing methodologies struggle to effectively validate the behavior of conversational AI systems under interactive and multi-granularity integration scenarios. To this end, it proposes the first hierarchical testing framework specifically designed for conversational AI, systematically covering verification across multiple levelsβfrom language and AI components, through single-agent systems, to multi-agent configurations. The approach integrates principles from software testing theory, architectural analysis of dialogue systems, and AI behavior validation techniques to construct a layered testing strategy. Experimental results demonstrate that the framework significantly enhances system reliability and consistency across different integration layers, offering a scalable and structured testing paradigm for complex conversational AI systems.
π Abstract
Conversational AI systems combine AI-based solutions with the flexibility of conversational interfaces. However, most existing testing solutions do not straightforwardly adapt to the characteristics of conversational interaction or to the behavior of AI components. To address this limitation, this Ph.D. thesis investigates a new family of testing approaches for conversational AI systems, focusing on the validation of their constituent elements at different levels of granularity, from the integration between the language and the AI components, to individual conversational agents, up to multi-agent implementations of conversational AI systems