Score
Design, implement, and analyze evaluation protocols, experiments, and quantitative/qualitative metrics that measure the performance and interaction dynamics of human–AI teams, covering human-AI teaming, HAT assessments, and LLM-mediated teaming. This includes operationalizing and measuring task efficiency and rewards, changes in planning and decision behavior, expertise-moderated interaction effects, and statistical comparisons across different guidance or interface conditions.
Current human-AI collaboration (HAIC) evaluation lacks a unified framework capable of accommodating the heterogeneity and dynamic reciprocity inherent across AI-centered, human-centered, and symbiotic HAIC paradigms. Method: This paper proposes the first structured evaluation framework tailored to all three HAIC modes, introducing a novel multi-dimensional evaluation decision tree that integrates quantitative and qualitative metrics—combining objective and subjective dimensions—and formally modeling dynamic reciprocity throughout the collaborative process. The framework is validated empirically across four domains—manufacturing, healthcare, finance, and education—via systematic literature review and cross-domain adaptation. Contribution/Results: Results demonstrate significant improvements in evaluation specificity, interpretability, and practical guidance value. The framework establishes a methodological foundation for scientifically measuring HAIC effectiveness, enabling rigorous, context-sensitive assessment of collaborative outcomes.
Current research on human–autonomy teams (HATs) is highly fragmented, focusing narrowly on isolated phases or singular challenges—such as trust calibration—without a systemic understanding of long-term adaptability. Method: Adopting a process-dynamics perspective, this study employs the T⁴ framework (Team Formation, Task and Role Development, Team Evolution, Team Optimization) and integrates systematic literature review (SLR) with cross-phase collaborative assessment modeling to achieve the first holistic, dynamic integration of HAT research across the entire lifecycle. Contribution/Results: We propose a “task–team bidirectional adaptation” analytical paradigm, uncovering core adaptive mechanisms—including role allocation, shared mental models, and backup behaviors. Six critical mechanisms influencing long-term collaborative efficacy are identified, and a comprehensive HAT adaptability assessment framework spanning the full lifecycle is established. This work provides both a theoretical foundation and an actionable roadmap for designing adaptive human–autonomy teams.
Current AI evaluation practices overly emphasize model accuracy while neglecting calibration and safety in human-AI collaboration, often leading to misuse or underestimation of system capabilities. This work proposes a novel evaluation framework centered on “team readiness,” introducing a four-dimensional metric system that quantifies outcomes, dependency behaviors, safety signals, and learning evolution directly from real collaborative interactions. The framework integrates interaction trajectory analysis, behavioral calibration measurement, error recovery assessment, and governance indicators, aligning with the Understand–Control–Improve (U-C-I) lifecycle of human-AI teamwork. By doing so, it enables comparable and reproducible evaluations of calibration quality, error recovery capacity, and governance maturity, thereby advancing safer and more accountable research in human-AI collaboration.
Existing research platforms predominantly focus on simplified tasks or single-perspective analysis, limiting their capacity to support interdisciplinary empirical studies of complex human-AI collaborative decision-making. To address this gap, we introduce CREW—an open-source platform featuring a novel modular architecture designed for ecologically valid collaborative scenarios. CREW integrates cognitive experimental paradigms, real-time multimodal physiological signal acquisition (EEG, ECG, EMG), human-guided reinforcement learning benchmarks (PPO, SAC), and an extensible task framework—unifying human behavioral modeling and AI algorithm evaluation. CREW is the first tool enabling real-time, multidisciplinary, high-ecological-validity studies of human-AI teams. Within one week, we validated the platform across 50 participants: results demonstrate significant improvements in task flexibility, depth of human engagement, and rigor of algorithmic assessment. CREW thus establishes a unified experimental foundation for both foundational research and applied development in human-AI collaborative decision-making.
This study addresses critical safety, fairness, and controllability challenges in Large Foundation Models (LFMs) and Human-AI (HAI) collaboration, aiming to establish trustworthy, socially beneficial human–machine partnerships. Methodologically, it introduces the first LPtM-driven four-dimensional analytical framework—comprising capability augmentation, model feedback, team-level coordination, and ethical governance—and proposes a theory of staged advancement in collaborative intelligence. Integrating techno-sociological analysis, interdisciplinary bibliometrics, case-based deconstruction, and multi-stakeholder governance modeling, the research systematically identifies 12 key collaborative patterns, seven cross-cutting challenges, and distills five actionable policy recommendations. The findings constitute the first comprehensive, empirically grounded benchmark for global HAI system design and governance, offering both conceptual rigor and practical applicability for advancing responsible AI integration in socio-technical systems.
This study addresses the lack of industry-compliant evaluation methodologies in existing AI research for air traffic control (ATC) tasks, which often fail to reflect real-world operational environments. To bridge this gap, the work introduces— for the first time—the legally mandated ATC training assessment framework into AI agent testing. It proposes a human-in-the-loop evaluation paradigm grounded in regulatory-certified simulator curricula, wherein domain-expert instructors conduct contextually accurate assessments of AI agent performance. This approach aligns AI capabilities with established human professional standards, substantially narrowing the divide between academic research and actual ATC operations, and lays a foundational framework for future human-AI collaborative air traffic management systems.
This study investigates how large language models (LLMs) influence attention allocation, decision-making behavior, and situational awareness among individuals with varying expertise during search-and-rescue tasks. Using a simulated environment, the research integrates eye-tracking data, behavioral logs, and task performance metrics to compare expert and novice performance with and without LLM guidance. Results indicate that while LLM assistance improves unit efficiency—yielding higher rewards and more rescues per action—it does not increase total rescues. Eye-tracking reveals a shift in user attention toward the chat interface; experts maintain environmental scanning and actively cross-verify AI suggestions, whereas novices tend to rely passively on model outputs. The study proposes a “verification loop” mechanism, underscoring the critical role of real-time alignment between AI recommendations and environmental states in preserving situational awareness during human-AI collaboration.
This study addresses the challenge of predicting collaborative team performance and enabling timely interventions without relying on task-specific prior knowledge. The authors propose TRIBE, a method that models early communication behaviors to cluster teams into predictive “behavioral tribes,” achieving effective performance forecasting using only initial-stage data (e.g., the first 10% of task duration). TRIBE is the first approach to enable cross-task, task-agnostic team performance prediction, revealing distinct influences of AI agents versus human advisors on team behavioral trajectories and demonstrating that successful teams maintain behavioral flexibility throughout collaboration. Experimental results across four heterogeneous datasets show that TRIBE significantly outperforms baseline methods, offering improvements in both prediction accuracy and inference speed.
This study addresses the lack of infrastructure supporting reproducible, longitudinal, and real-time human–AI collaboration experiments, which has hindered systematic investigation into how design attributes of AI teammates influence team trust, coordination, and decision-making. To bridge this gap, the authors introduce TRAIL, a novel platform that embeds configurable and reproducible AI teammates within an instrumented, authentic collaborative environment, enabling longitudinal experimentation and behavioral analysis. TRAIL innovatively integrates the Big Five personality model, selective messaging channels, a dual-memory architecture, chained experimental scheduling, and textual similarity analysis tools to systematically modulate AI personality, communication timing, and interaction style. In a six-round classroom study with 51 students, TRAIL sustained stable AI collaboration and revealed significant differential effects of AI personality on team perceptions of contribution, linguistic alignment, group atmosphere, and reliance on the AI teammate.
This study addresses the lack of reliable instruments for assessing subjective collaboration quality in human–AI teamwork. Drawing on joint activity theory and evolutionary cooperation theory, it develops and validates two novel, cross-agent and cross-context subjective scales: the Perceived Cooperativeness Scale (PCS) and the Teamness Perception Scale (TPS), which respectively measure perceived cooperativeness and sense of teaming. Through rigorous psychometric methods, the factorial structure, reliability, and construct validity of both scales were established across three studies (N = 409) and diverse experimental scenarios—including card games, interactions with large language models, and decision support systems. Findings demonstrate that the scales effectively differentiate partners exhibiting varying levels of collaborative quality, offering a unified and comparable tool for quantifying both human–AI and human–human collaboration.