Score
The practice of measuring and quantifying users' mental effort and information-processing burden using validated metrics and observational/experimental methods; used to evaluate how interventions, explanations, or tools affect comprehension, attitudes, and well‑being. It includes designing tasks and instruments to compare cognitive load across conditions and to fold load into multi-dimensional decision metrics.
Assessing cognitive readiness via wearable devices in real-world settings faces challenges including high inter-individual variability, complex coupling among multimodal physiological signals (e.g., HRV, EDA, EEG, eye-tracking), and stringent hardware constraints. To address these, this study proposes the first physiological biomarker classification framework specifically designed for *in-the-wild* deployment. It establishes a standardized measurement taxonomy that balances personalization with cross-device compatibility. We systematically review state-of-the-art acquisition and interpretation methods for key physiological modalities, integrating human-computer interaction (HCI) and ubiquitous computing analytical paradigms. The resulting comprehensive practice guideline explicitly addresses three critical dimensions: (1) conceptual assessment scope, (2) hardware interoperability across consumer-grade wearables, and (3) ecological validity under naturalistic conditions. This work delivers a deployable foundation for real-time cognitive monitoring, with direct applicability to education, defense, and human factors engineering domains.
This study addresses how time pressure and item difficulty in web-based questionnaires can induce stress and cognitive overload, thereby compromising data quality and user experience. Traditional low-frequency self-report methods are insufficient for capturing dynamic state changes during brief tasks. To overcome this limitation, the authors conducted a 2×2 within-subjects experiment manipulating these factors while simultaneously collecting multimodal physiological and behavioral data—including eye-tracking, electrocardiography (ECG), electrodermal activity (EDA), and mouse trajectories. By integrating statistical modeling with machine learning techniques, they identified distinct, short-latency patterns of physiological and behavioral responses that reliably differentiate cognitive-affective states. The findings demonstrate the feasibility of real-time detection of such states in digital environments and provide critical technical foundations for developing adaptive, user-aware online questionnaire systems.
Prior research lacks systematic comparison between subjective self-reports and neurophysiological measures of cognitive load (CL) in data visualization evaluation. Method: We conducted an experiment integrating visualization literacy assessment and spatial visualization tasks, simultaneously recording 32-channel EEG signals. A graph attention network (GAT) was developed to estimate mental workload (MW) from EEG, and results were compared against established subjective scales (e.g., NASA-TLX). Contribution/Results: Significant discrepancies emerged between EEG-derived MW estimates and subjective ratings across task difficulty levels. EEG proved sensitive to unconscious cognitive effort—unreported in self-assessments—and revealed dynamic CL fluctuations invisible to introspection. This study provides the first empirical validation in visualization research that neurophysiological metrics meaningfully complement subjective evaluation. It establishes a novel, objective, and fine-grained paradigm for CL assessment, offering both methodological innovation and theoretical grounding for advancing visualization usability evaluation.
This study quantifies students’ cognitive effort during educational gaming to enable dynamic adaptation of instructional materials. Using functional near-infrared spectroscopy (fNIRS), we measured prefrontal cortical oxygenated hemoglobin signals, extracted temporal statistical and functional connectivity features, and applied multi-model machine learning to predict behavioral performance (accuracy: 58–67%). Our key contribution is the novel dual-metric framework—“relative neural efficiency” and “relative neural engagement”—which jointly integrates neural activation magnitude with behavioral output to robustly characterize trends in cognitive effort. Although single-trial prediction accuracy remains modest, both metrics exhibit strong consistency with task difficulty and learning phase progression. Results demonstrate that this paradigm enables objective, continuous monitoring of cognitive load dynamics. It thus provides an interpretable, deployable neurophysiological foundation for adaptive educational systems grounded in real-time neural feedback.
Quantifying cognitive load during conceptual design—particularly when designers envision products with relatively moving components—remains challenging due to the lack of interpretable, task-sensitive neural metrics. Method: This study proposes and validates inter-band relative power difference (inter-BRPD), a novel EEG-derived metric designed to efficiently and interpretably characterize mental effort in motion exploration tasks (MET) and concept generation tasks (CGT). Leveraging MNE-Python for preprocessing, inter-BRPD achieves high reliability and validity using approximately 50% fewer parameters than conventional EEG indices. Results: Inter-BRPD significantly discriminates between distinct task load levels (p < 0.01) and demonstrates excellent test–retest reliability (ICC = 0.92). It provides a neurocognitively grounded, computationally lightweight framework for analyzing conceptual design processes and empirically validating design intervention tools.
A critical challenge in explainable AI (XAI) and human-AI collaboration is determining *when* to provide explanations—i.e., real-time identification of genuine explanation needs—yet existing approaches rely on static, subjective assumptions and fail to dynamically capture users’ contextual demands. Method: We propose the first holistic, real-time explanation-need recognition framework integrating user behavior, system events, and physiological-emotional signals. Through systematic literature synthesis and empirical validation, we identify and verify 39 measurable, generalizable, and triggerable user-side indicators. We organize them into a cross-dimensional taxonomy (behavioral, system-event, and affective/physiological) and develop a demand-type mapping model. Contribution/Results: Grounded in online experiments, self-reports, and qualitative coding, we establish a structured indicator catalog comprising 17 behavioral, 8 system-event, and 14 affective/physiological measures, and design its runtime telemetry integration. The framework enables dynamic, precise, and temporally appropriate automated explanation triggering, validated in both prototype and production environments.
This study addresses a critical gap in understanding the unique psychological and physiological burdens arising from prolonged AI tool use in academic settings. Drawing on grounded theory, it systematically codes and thematically analyzes open-ended responses from 1,054 Filipino university students to propose “AI fatigue” as a distinct construct. The research identifies five core dimensions—cognitive overload, motivational disengagement, moral unease, physical tension, and attentional drift—each operationalized through two empirically derived indicators. Furthermore, it develops a staged accumulation model that elucidates the dynamic interplay and temporal evolution of these multidimensional stressors. By delineating the structural and processual characteristics of AI fatigue, this work establishes a foundational theoretical framework for future scale development and cross-contextual investigations.
This study addresses the limitations of existing digital mental health tools, which often rely on rigid scripts and struggle to support users in naturally articulating and cognitively reappraising stressful events. The authors developed and evaluated a single-session AI intervention powered by GPT-4o that guides working professionals through structured dialogues to reframe workplace stressors. A multimodal assessment framework—integrating a RoBERTa-based sentiment classifier, an LLM-derived stress scorer, and thematic analysis—was employed to evaluate outcomes. Deployed in a real-world workplace setting, this work demonstrates for the first time the feasibility of LLM-driven cognitive reappraisal, significantly reducing perceived stress intensity and improving coping mindsets. Automated analyses revealed a progressive decline in negative affect throughout the dialogue, and participants acknowledged the value of guided reflection, though they also noted the interaction’s scripted feel and excessive length, highlighting tensions between AI-emulated empathy and interaction design.
This study addresses the absence of an objective and reproducible evaluation framework for artificial general intelligence (AGI), which has led to subjective assessments of progress and challenges in governance. Drawing on foundations from psychology, neuroscience, and cognitive science, this work proposes a systematic cognitive taxonomy comprising ten core capabilities grounded in human cognition. It introduces targeted retention tasks designed to evaluate system performance across these dimensions, thereby generating multidimensional cognitive profiles. The approach enables an operational decomposition and empirical measurement of AGI progress, offering an initial benchmark to identify system strengths and weaknesses and to advance the standardization and transparency of AGI research.
Current measures of AI reliance primarily rely on output adoption or subjective self-reports, which inadequately capture the allocation of cognitive effort between users and AI during task execution. This work proposes a counterfactual workflow-based simulation method that models the steps users would take without AI assistance to quantify the proportion of cognitive effort offloaded to the AI. Introducing a novel metric—the Offloading Score—it provides a more precise measure of AI dependence. This score effectively captures dynamic shifts in reliance under time pressure, facilitating both user self-reflection and system-level interventions. In a programming study with 40 developers, the Offloading Score detected a statistically significant 43% increase in reliance under time pressure (p = 0.018), outperforming conventional metrics and revealing that heightened dependence manifests as increased delegation of subtasks and direct reuse of AI-generated outputs.
Continuously estimating the dynamic changes in human cognitive capacity remains challenging. This work proposes a theory-driven multimodal learning framework that models cognitive capacity as a two-dimensional physiological state space defined by mental effort and stress. A dual-stream neural network encodes heart rate variability (HRV) and electrodermal activity (EDA) signals separately, which are then integrated via a late fusion strategy coupled with task-specific probabilistic output heads to jointly predict both dimensions. By grounding the two-dimensional physiological representation in established cognitive theories, the approach effectively discriminates between states such as efficient engagement and overload-induced stress. Evaluated on the SWELL-KW dataset, the model achieves balanced accuracies of 70.0% for stress and 72.2% for effort, demonstrating the efficacy of theory-guided supervision and multimodal fusion while sensitively capturing task-induced dynamics in cognitive demand.