Score
Designs, implements, and analyzes evaluation frameworks and empirical protocols for human-agent systems, including collaboration benchmarks, graph-based models of interaction, and human-in-the-loop system designs. Builds metrics and studies that measure task outcomes and recovery, assess clarification and feedback quality, quantify control calibration and safety, and compare interaction costs and initiative to evaluate system performance and user–agent collaboration.
Current research on human–autonomy teams (HATs) is highly fragmented, focusing narrowly on isolated phases or singular challenges—such as trust calibration—without a systemic understanding of long-term adaptability. Method: Adopting a process-dynamics perspective, this study employs the T⁴ framework (Team Formation, Task and Role Development, Team Evolution, Team Optimization) and integrates systematic literature review (SLR) with cross-phase collaborative assessment modeling to achieve the first holistic, dynamic integration of HAT research across the entire lifecycle. Contribution/Results: We propose a “task–team bidirectional adaptation” analytical paradigm, uncovering core adaptive mechanisms—including role allocation, shared mental models, and backup behaviors. Six critical mechanisms influencing long-term collaborative efficacy are identified, and a comprehensive HAT adaptability assessment framework spanning the full lifecycle is established. This work provides both a theoretical foundation and an actionable roadmap for designing adaptive human–autonomy teams.
Existing LLM-agent evaluation benchmarks predominantly assume full automation, neglecting realistic human-AI collaboration scenarios. Method: We propose PULSE—the first systematic evaluation framework explicitly designed for human-agent co-execution—integrating explicit user feedback with a satisfaction prediction model trained on over 15,000 real-world user interactions, augmented by pseudo-labeling and rigorously validated via A/B testing. Contribution/Results: Experiments reveal a significant negative correlation between scores on mainstream benchmarks and developers’ actual satisfaction; core architectural components—including LLM backbones, planning strategies, and memory mechanisms—are substantially underestimated in their impact on user experience by conventional metrics. Moreover, PULSE reduces confidence intervals of key evaluation metrics by 40%, establishing a more reliable, interpretable, and user-centered assessment paradigm for software agents.
Current AI evaluation predominantly relies on static responses, which inadequately capture the systematic capabilities of large language models in dynamic scenarios involving tool use, environmental interaction, and multi-agent collaboration. This work proposes “interactive evaluation” as a distinct paradigm, introducing a two-axis taxonomy to clarify its design principles and reporting standards. By leveraging trajectory modeling and a multi-dimensional scoring mechanism, the framework extracts evidence from interaction processes to holistically assess model performance across dimensions such as procedural fidelity, recoverability, coordination, robustness, and system-level efficacy. The study delineates core challenges in interactive evaluation and redefines the logical mapping from empirical evidence to performance judgments, thereby establishing a theoretical foundation for a unified, comparable, and interpretable next-generation AI evaluation framework.
Current AI evaluation practices overly emphasize model accuracy while neglecting calibration and safety in human-AI collaboration, often leading to misuse or underestimation of system capabilities. This work proposes a novel evaluation framework centered on “team readiness,” introducing a four-dimensional metric system that quantifies outcomes, dependency behaviors, safety signals, and learning evolution directly from real collaborative interactions. The framework integrates interaction trajectory analysis, behavioral calibration measurement, error recovery assessment, and governance indicators, aligning with the Understand–Control–Improve (U-C-I) lifecycle of human-AI teamwork. By doing so, it enables comparable and reproducible evaluations of calibration quality, error recovery capacity, and governance maturity, thereby advancing safer and more accountable research in human-AI collaboration.
This study addresses the lack of industry-compliant evaluation methodologies in existing AI research for air traffic control (ATC) tasks, which often fail to reflect real-world operational environments. To bridge this gap, the work introduces— for the first time—the legally mandated ATC training assessment framework into AI agent testing. It proposes a human-in-the-loop evaluation paradigm grounded in regulatory-certified simulator curricula, wherein domain-expert instructors conduct contextually accurate assessments of AI agent performance. This approach aligns AI capabilities with established human professional standards, substantially narrowing the divide between academic research and actual ATC operations, and lays a foundational framework for future human-AI collaborative air traffic management systems.
Existing AI evaluations predominantly rely on static benchmarking, failing to detect emergent harms—such as emotional dependency, social manipulation, and cognitive overload—that arise during prolonged human-AI interaction. To address this gap, we propose a novel dynamic evaluation paradigm centered on *interactional ethics*, moving beyond conventional single-turn output assessment. We introduce the first integrated framework for interactional ethical evaluation, combining controlled human-AI interaction experiments, NLP-based behavioral analysis, validated social science psychometric scales, and computational modeling of human impact. Our framework explicitly incorporates the evolution of human-AI relationships, downstream societal effects, and cognitive consequences as core evaluation dimensions. The project yields actionable principles for identifying interactional harms and scenario-specific evaluation protocols, thereby bridging a critical gap in safety assessment for generative AI deployed in authentic, long-term interaction settings. This work establishes a methodological foundation for context-aware, usage-oriented AI governance.
This work addresses a critical gap in existing evaluation frameworks, which often overlook the active role of humans in large language model (LLM)-driven human-AI collaboration systems. The authors propose the HAS-Framework, which uniquely models both humans and LLM agents as first-class participants with explicit roles, permissions, and communication pathways. Complementing this, they introduce HAS-Bench, a configurable benchmark that supports diverse modes of human involvement. Collaboration dynamics are formalized through a graph-based representation, and system performance is assessed via multidimensional metrics—including clarification quality, feedback utilization, and control calibration—capturing both process-level interactions and outcome-level efficacy. Experiments across six domains demonstrate that strategically configuring the timing, modality, and role of human participation significantly enhances task completion rates and system robustness.
This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.
While autonomous software agents have enhanced development efficiency, their errors and novel failure modes necessitate effective human oversight—an area lacking empirical investigation into how developers actually conduct such supervision. This study addresses this gap through semi-structured interviews with 17 experienced developers, integrating theories from human–AI collaboration and software engineering. It identifies four emergent forms of supervision: proactive control, collaborative planning, real-time monitoring, and post-hoc review—challenging the traditional view of supervision as merely reactive and revealing its inherently preventive nature. The work further distills practical heuristics, including verifying code correctness through test outcomes, offering crucial empirical insights and design implications for human-centered agent development and software engineering practice.
Existing agent evaluation methods rely on static benchmarks, which struggle to capture realistic failure modes in multi-step dynamic interactions and lack effective metrics for assessing interaction quality and coverage. To address these limitations, this work proposes VISTA—the first hybrid user simulation framework that integrates both UI and API interactions—and introduces a six-dimensional evaluation metric suite to comprehensively measure interaction realism, capability coverage, and effectiveness. Experiments in e-commerce shopping and educational customer service scenarios demonstrate that VISTA significantly enhances the realism and comprehensiveness of agent evaluations, generating assessment outcomes that are both more authentic and more broadly representative than those produced by current approaches.
As the scale of agent skill repositories grows, the absence of systematic evaluation mechanisms undermines guarantees regarding skill utility, quality, and safety. This work introduces a unified framework for skill evaluation and evolution by formally categorizing skill evolution into four paradigms: execution feedback, trajectory distillation, compression, and reinforcement learning. Through comprehensive multidimensional benchmarking, the study systematically analyzes six existing evaluation methodologies, uncovering their structural gaps and metric limitations. By shifting the paradigm from isolated skill construction to evaluation-driven automated evolution, this research lays the foundation for developing general-purpose, efficient, and verifiably safe skill ecosystems.