evaluate human-agent systems

Designs, implements, and analyzes evaluation frameworks and empirical protocols for human-agent systems, including collaboration benchmarks, graph-based models of interaction, and human-in-the-loop system designs. Builds metrics and studies that measure task outcomes and recovery, assess clarification and feedback quality, quantify control calibration and safety, and compare interaction costs and initiative to evaluate system performance and user–agent collaboration.

evaluatehuman-agentsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

How can we assess human-agent interactions? Case studies in software agent design

Oct 10, 2025
VC
Valerie Chen
🏛️ Carnegie Mellon University | All Hands AI

Existing LLM-agent evaluation benchmarks predominantly assume full automation, neglecting realistic human-AI collaboration scenarios. Method: We propose PULSE—the first systematic evaluation framework explicitly designed for human-agent co-execution—integrating explicit user feedback with a satisfaction prediction model trained on over 15,000 real-world user interactions, augmented by pseudo-labeling and rigorously validated via A/B testing. Contribution/Results: Experiments reveal a significant negative correlation between scores on mainstream benchmarks and developers’ actual satisfaction; core architectural components—including LLM backbones, planning strategies, and memory mechanisms—are substantially underestimated in their impact on user experience by conventional metrics. Moreover, PULSE reduces confidence intervals of key evaluation metrics by 40%, establishing a more reliable, interpretable, and user-centered assessment paradigm for software agents.

Addressing limitations of automated benchmarks for real-world usageAssessing impact of LLM design choices on user satisfaction ratesEvaluating human-agent interaction in collaborative software systems

Current AI evaluation predominantly relies on static responses, which inadequately capture the systematic capabilities of large language models in dynamic scenarios involving tool use, environmental interaction, and multi-agent collaboration. This work proposes “interactive evaluation” as a distinct paradigm, introducing a two-axis taxonomy to clarify its design principles and reporting standards. By leveraging trajectory modeling and a multi-dimensional scoring mechanism, the framework extracts evidence from interaction processes to holistically assess model performance across dimensions such as procedural fidelity, recoverability, coordination, robustness, and system-level efficacy. The study delineates core challenges in interactive evaluation and redefines the logical mapping from empirical evidence to performance judgments, thereby establishing a theoretical foundation for a unified, comparable, and interpretable next-generation AI evaluation framework.

agent benchmarksevaluation paradigmsinteractive evaluation

Current AI evaluation practices overly emphasize model accuracy while neglecting calibration and safety in human-AI collaboration, often leading to misuse or underestimation of system capabilities. This work proposes a novel evaluation framework centered on “team readiness,” introducing a four-dimensional metric system that quantifies outcomes, dependency behaviors, safety signals, and learning evolution directly from real collaborative interactions. The framework integrates interaction trajectory analysis, behavioral calibration measurement, error recovery assessment, and governance indicators, aligning with the Understand–Control–Improve (U-C-I) lifecycle of human-AI teamwork. By doing so, it enables comparable and reproducible evaluations of calibration quality, error recovery capacity, and governance maturity, thereby advancing safer and more accountable research in human-AI collaboration.

decision-makingevaluation metricshuman-AI collaboration

This study addresses the lack of industry-compliant evaluation methodologies in existing AI research for air traffic control (ATC) tasks, which often fail to reflect real-world operational environments. To bridge this gap, the work introduces— for the first time—the legally mandated ATC training assessment framework into AI agent testing. It proposes a human-in-the-loop evaluation paradigm grounded in regulatory-certified simulator curricula, wherein domain-expert instructors conduct contextually accurate assessments of AI agent performance. This approach aligns AI capabilities with established human professional standards, substantially narrowing the divide between academic research and actual ATC operations, and lays a foundational framework for future human-AI collaborative air traffic management systems.

AI EvaluationAir Traffic ControlHuman-in-the-Loop

Towards interactive evaluations for interaction harms in human-AI systems

May 17, 2024
LI
Lujain Ibrahim
🏛️ University of Oxford | OpenAI

Existing AI evaluations predominantly rely on static benchmarking, failing to detect emergent harms—such as emotional dependency, social manipulation, and cognitive overload—that arise during prolonged human-AI interaction. To address this gap, we propose a novel dynamic evaluation paradigm centered on *interactional ethics*, moving beyond conventional single-turn output assessment. We introduce the first integrated framework for interactional ethical evaluation, combining controlled human-AI interaction experiments, NLP-based behavioral analysis, validated social science psychometric scales, and computational modeling of human impact. Our framework explicitly incorporates the evolution of human-AI relationships, downstream societal effects, and cognitive consequences as core evaluation dimensions. The project yields actionable principles for identifying interactional harms and scenario-specific evaluation protocols, thereby bridging a critical gap in safety assessment for generative AI deployed in authentic, long-term interaction settings. This work establishes a methodological foundation for context-aware, usage-oriented AI governance.

Address risks like manipulation and overreliance in repeated interactionsInteractive AI systems need ethics-focused evaluation methodsStatic AI tests miss harms from human-AI interaction

Latest Papers

What's happening recently
View more

This work addresses a critical gap in existing evaluation frameworks, which often overlook the active role of humans in large language model (LLM)-driven human-AI collaboration systems. The authors propose the HAS-Framework, which uniquely models both humans and LLM agents as first-class participants with explicit roles, permissions, and communication pathways. Complementing this, they introduce HAS-Bench, a configurable benchmark that supports diverse modes of human involvement. Collaboration dynamics are formalized through a graph-based representation, and system performance is assessed via multidimensional metrics—including clarification quality, feedback utilization, and control calibration—capturing both process-level interactions and outcome-level efficacy. Experiments across six domains demonstrate that strategically configuring the timing, modality, and role of human participation significantly enhances task completion rates and system robustness.

Collaboration EvaluationConfigurable InteractionHuman Participation

This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.

actionable feedbackagent-human interactioncrowdsourced testing

While autonomous software agents have enhanced development efficiency, their errors and novel failure modes necessitate effective human oversight—an area lacking empirical investigation into how developers actually conduct such supervision. This study addresses this gap through semi-structured interviews with 17 experienced developers, integrating theories from human–AI collaboration and software engineering. It identifies four emergent forms of supervision: proactive control, collaborative planning, real-time monitoring, and post-hoc review—challenging the traditional view of supervision as merely reactive and revealing its inherently preventive nature. The work further distills practical heuristics, including verifying code correctness through test outcomes, offering crucial empirical insights and design implications for human-centered agent development and software engineering practice.

autonomous systemsdeveloper productivityempirical study

Existing agent evaluation methods rely on static benchmarks, which struggle to capture realistic failure modes in multi-step dynamic interactions and lack effective metrics for assessing interaction quality and coverage. To address these limitations, this work proposes VISTA—the first hybrid user simulation framework that integrates both UI and API interactions—and introduces a six-dimensional evaluation metric suite to comprehensively measure interaction realism, capability coverage, and effectiveness. Experiments in e-commerce shopping and educational customer service scenarios demonstrate that VISTA significantly enhances the realism and comprehensiveness of agent evaluations, generating assessment outcomes that are both more authentic and more broadly representative than those produced by current approaches.

failure modesinteractive agent evaluationrealistic user behaviors

As the scale of agent skill repositories grows, the absence of systematic evaluation mechanisms undermines guarantees regarding skill utility, quality, and safety. This work introduces a unified framework for skill evaluation and evolution by formally categorizing skill evolution into four paradigms: execution feedback, trajectory distillation, compression, and reinforcement learning. Through comprehensive multidimensional benchmarking, the study systematically analyzes six existing evaluation methodologies, uncovering their structural gaps and metric limitations. By shifting the paradigm from isolated skill construction to evaluation-driven automated evolution, this research lays the foundation for developing general-purpose, efficient, and verifiably safe skill ecosystems.

agent skill evaluationagentic systemsbenchmarking

Hot Scholars

AW

April Wang

ETH Zurich
Human Computer Interaction
JC

Jiaju Chen

Computer Science, Rice University
Human-Computer InteractionNatural Language Processing
TH

Tilo Hartmann

Vrije Universiteit (VU) Amsterdam
Communication ScienceMedia PsychologyVirtual RealityVideo Games
AE

Ali Eslami

Associate Professor of EECS, Wichita State University
Cyber-Physical SystemsError Correction CodingQuantum ComputingInternet of Things