🤖 AI Summary
This study addresses the challenge in XR sports viewing where voice queries often omit critical information—such as player, time, location, or metric—leading to implicit parsing errors that users struggle to detect and correct. To tackle this, the work proposes a taxonomy of ambiguity tailored for XR voice interaction and an externalization design space that treats transparency and correctability as orthogonal dimensions. By integrating spatialized visual cues with auxiliary analytical views, the system externalizes parsed query results in real time within the immersive XR interface. A user study demonstrates that this approach significantly enhances users’ ability to inspect and understand ambiguities, encouraging more explicit verbal corrections; however, only 38% of misinterpretations were actually corrected, with effectiveness varying across ambiguity types.
📝 Abstract
XR sports viewing enables spectators to follow play from immersive, spatially anchored perspectives while accessing contextual analytics directly within the scene. In such settings, speech offers a practical interaction modality because text entry and menu navigation can interrupt attention during fast-paced gameplay. However, spoken queries are often underspecified: viewers may omit which player, time period, field location, or metric they intend. When systems resolve these ambiguities implicitly, their assumptions remain hidden, making misinterpretations difficult to notice and correct (repair). We investigate how externalizing a system's interpretation of spoken queries can support inspection and correction of such misunderstandings in XR sports viewing. Through a formative study, we identified four recurring ambiguity types (referential, spatial, temporal, and metric) that characterize ambiguous spoken queries in this context. We develop a design space that organizes externalization along three dimensions (ambiguity type, interpretation state, externalization strategy) and instantiate it in an interactive XR soccer viewing system that combines situated visual cues with supporting analytic views. A within-subjects user study (N=16) comparing externalized interpretation against a voice-only baseline reveals that externalization is associated with higher inspectability on most measured dimensions and increased explicit repair language overall. However, repair occurred in only 38% of misaligned externalization trials, and this visibility-action gap varied by ambiguity type, indicating that transparency and correction affordance are orthogonal design axes.