π€ AI Summary
This study addresses the challenge of proactively determining both the timing and content of interventions for voice assistants operating on first-person video streams. To this end, it proposes HoloAssist, a proactive voice guidance framework that, for the first time, jointly optimizes when to intervene and how to instruct. The approach leverages source separation, speech resynthesis, and fine-tuning of an omni-modal large language model to enable real-time perception and generation. Furthermore, Direct Preference Optimization (DPO) is introduced to align the modelβs proactive interaction behavior with human preferences. Experimental results demonstrate that the proposed method significantly outperforms zero-shot baselines in both intervention timing accuracy and content relevance, while also achieving higher human preference ratings.
π Abstract
Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance. We fine-tune an omni-modal LLM with our data, and further improve its proactive intervention behavior with direct preference optimization. Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.