You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation methods struggle to effectively assess the long-term performance and evolving user trust of AI-driven dynamic head-up augmented reality (AR) interfaces in multitasking environments with continuously shifting contexts. This work proposes a dynamic evaluation paradigm grounded in three key shifts: moving from isolated metrics to performance–operational characteristics (POC) interference frontiers, from static conditions to context-trajectory-based assessment, and from single-session studies to longitudinal trust tracking. Integrating optical see-through head-mounted displays (OST-HMDs), multitasking scenarios, context-trajectory sampling, and decoupled cognitive load analysis, the study establishes the first multidimensional, quantifiable evaluation framework tailored for dynamic AR interfaces. This approach exposes the limitations of conventional snapshot evaluations and provides an empirical foundation for systematic interface optimization.
📝 Abstract
AI-mediated heads-up augmented reality (AR) replaces fixed interfaces with dynamically adapting ones that decide what information to present, in what form, and when, based on a continually changing context that cannot be fully anticipated beforehand. Although it remains an interface, its behavior over time is only partially specified at design time. We argue that this shift requires a corresponding change in evaluation: from snapshots to trajectories. A fixed interface is evaluated in a snapshot --- one context, one session, one set of task-performance metrics. A fluid interface must be evaluated over a trajectory --- a sequence of contexts with transitions, sampled from the distribution the interface will actually encounter, and tracked long enough for user trust to form, evolve, and potentially deteriorate. Drawing on the literature for heads-up AR multitasking enabled by optical see-through head-mounted displays (OST-HMDs), we find that current evaluation practice remains largely snapshot-based. Most studies use fixed-condition, single-session designs; interference between concurrent tasks is rarely quantified directly; and commonly used workload measures cannot disentangle cognitive load attributable to individual tasks. To address these limitations, we argue for three shifts: from isolated metrics to Performance Operating Characteristic (POC) interference frontiers, from fixed conditions to evaluation over context trajectories, and from single-session snapshots to longitudinal trust measurement.
Problem

Research questions and friction points this paper is trying to address.

AI-mediated AR
multitasking evaluation
context trajectories
user trust
cognitive load
Innovation

Methods, ideas, or system contributions that make the work stand out.

Performance Operating Characteristic
context trajectories
longitudinal trust measurement
multitasking evaluation
AI-mediated AR
🔎 Similar Papers
No similar papers found.