VoCa: Designing Speech-Canvas Interaction for Voice-Based Conversational Agents
This study addresses the lack of collaborative visual canvas interaction in voice-based agents by proposing VoCa, a system that enables dynamic coordination between auditory and visual information. Methodologically, we construct a voice-canvas interaction design space and employ user observations, design workshops, and prototype experiments to align spoken dialogue with canvas object creation, annotation, and attention guidance, thereby optimizing multi-turn interaction experiences. The research validates the interactive potential of cross-modal collaboration and reveals core challenges in coordinating “saying” and “showing.” Ultimately, this work provides both theoretical foundations and practical guidelines for interface design and interaction paradigms in multimodal intelligent agents.