π€ AI Summary
This study addresses the lack of collaborative visual canvas interaction in voice-based agents by proposing VoCa, a system that enables dynamic coordination between auditory and visual information. Methodologically, we construct a voice-canvas interaction design space and employ user observations, design workshops, and prototype experiments to align spoken dialogue with canvas object creation, annotation, and attention guidance, thereby optimizing multi-turn interaction experiences. The research validates the interactive potential of cross-modal collaboration and reveals core challenges in coordinating βsayingβ and βshowing.β Ultimately, this work provides both theoretical foundations and practical guidelines for interface design and interaction paradigms in multimodal intelligent agents.
π Abstract
People write and sketch while speaking to explain, organize, and develop content together. Inspired by these practices, we investigate how voice agents can use a canvas alongside speech in multi-turn conversations with users. We conducted a two-part formative study: an observational study of how pairs coordinated speech and boardwork, followed by a design workshop that informed a design space for speech-canvas interaction with voice agents. Building on these insights, we developed VoCa, a voice agent that coordinates speech with visual object creation, annotation, and attention guidance. A five-day deployment with 18 participants examined usability, experiences of speech-canvas interaction, patterns of use, and desired improvements. Participants'experiences highlighted opportunities for speech-canvas interaction in learning, work, and daily life, alongside challenges in coordinating what agents say and show in ways users can follow and influence. These findings inform how voice agents can use a canvas alongside speech in conversation.