Score
Designs and implements simulation systems that include live human participants in the control and feedback loop, enabling real-time human input, observation, and measurement while coordinating with simulated agents. Builds and operates experiment and analysis pipelines to run and scale HITL studies—managing multi‑agent scenarios, mediating human–agent interactions, capturing telemetry and user actions, and interfacing the simulation with external planners or controllers to evaluate behavior and performance.
This paper addresses the lack of a systematic taxonomy for human-AI interaction in agent-based modeling and simulation (ABMS) amid the rise of large language models (LLMs). We propose the first five-dimensional taxonomy—Why/When/What/Who/How—specifically designed for ABMS contexts. Integrating theories from human-computer interaction, ABMS methodology, and empirical LLM deployment practices, we conduct a structured literature review and cross-dimensional pattern analysis to synthesize existing human-AI interaction practices and identify recurrent interaction archetypes. The taxonomy fills a critical gap in systematic scholarly synthesis, clarifies paradigm boundaries and collaborative pathways for human-in-the-loop integration, and provides actionable design principles for ABMS tool development. Furthermore, it exposes current interaction blind spots—such as underexplored temporal dynamics and asymmetric agency configurations—thereby delineating concrete directions for future research on human-AI co-simulation and socio-technical alignment in ABMS.
This work addresses the limitations of existing human-in-the-loop (HITL) mechanisms in intelligent agent workflows, which are often tightly coupled with application logic, resulting in poor reusability, weak consistency, and limited scalability. To overcome these challenges, the paper proposes a decoupled HITL system architecture that abstracts human oversight into an independent component. By introducing explicit interfaces and a structured execution model, the approach cleanly separates human–machine interaction from business logic. Furthermore, it introduces a novel four-dimensional framework—comprising intervention conditions, role resolution, interaction semantics, and communication channels—to enable context-aware, controllable human intervention. This design achieves, for the first time, protocol-level reusability of HITL mechanisms, supporting consistent and scalable autonomy governance in multi-agent environments and laying a foundational infrastructure for system-level human–agent collaboration.
This study addresses the challenge non-technical researchers face in designing and analyzing complex experiments in multi-agent team dynamics. We propose VirTLab: an interactive 2D simulation platform powered by large language models (LLMs). Integrating team cognition theory with scalable agent modeling, VirTLab enables users—without programming expertise—to define environments, agent roles, tasks, and interaction rules, facilitating flexible simulation of coordination mechanisms, collective behavior, and emergent phenomena in human-AI collaboration. Its key contribution lies in balancing ecological validity and accessibility: spatialized agent behavior modeling, role-driven communication protocols, and real-time visualization empower both technical and non-technical researchers to conduct empirically grounded experiments. Evaluation demonstrates high fidelity between VirTLab’s simulated outputs and observed human team behavior, significantly lowering the barrier to entry for multi-agent experimentation.
Existing human-in-the-loop simulation approaches rely heavily on heuristic parameter tuning and lack data-driven personalization, resulting in insufficient fidelity. This work proposes a Real2Sim standardization pipeline that leverages user feedback on “safety and comfort” to identify a 12-dimensional set of individualized parameters at the pelvis-harness interface, using a six-degree-of-freedom viscoelastic model optimized via the CMA-ES algorithm. Intra-class correlation analysis distinguishes universal from subject-specific parameters, while a reproducible operating point eliminates ambiguity in harness tension. Remarkably, only five parameters require calibration to adapt the model to a new user. The calibrated model accurately reproduces real-world interaction envelopes and elicits biomechanically plausible gait adaptations, significantly enhancing simulation fidelity and enabling preclinical validation of personalized controllers.
This study addresses the lack of industry-compliant evaluation methodologies in existing AI research for air traffic control (ATC) tasks, which often fail to reflect real-world operational environments. To bridge this gap, the work introduces— for the first time—the legally mandated ATC training assessment framework into AI agent testing. It proposes a human-in-the-loop evaluation paradigm grounded in regulatory-certified simulator curricula, wherein domain-expert instructors conduct contextually accurate assessments of AI agent performance. This approach aligns AI capabilities with established human professional standards, substantially narrowing the divide between academic research and actual ATC operations, and lays a foundational framework for future human-AI collaborative air traffic management systems.
This study addresses the low automation level and high human dependency in scientific research workflows by proposing an Autonomous Simulation Agent (ASA) framework tailored for long-duration simulation tasks. Methodologically, the ASA integrates prompt engineering, automated code generation, remote high-performance computing (HPC) job scheduling, and multi-stage workflow orchestration, featuring a dynamically self-verifying architecture. A novel local-attention–global-supervision coordination mechanism enables 20 rounds of fully autonomous, human-free iteration. Evaluated on polymer chain conformational sampling, ASA-GPT-4o achieves near 100% task completion rate and sustains stable end-to-end operation across 20 consecutive cycles. The framework significantly enhances research efficiency, operational reliability, and experimental reproducibility, advancing the automation and robustness of computational science workflows.
This work addresses the complexity of human–robot interaction in multi-step, insertion-based, and fine teleoperation tasks by proposing a human-in-the-loop shared control framework that, for the first time, integrates diffusion policies into teleoperation systems. The approach combines human input with a point-cloud-based diffusion model to automatically adjust the end-effector orientation of a robotic arm, enabling high-dimensional manipulation through position-only control and thereby significantly simplifying the operator interface. Experimental results demonstrate that, compared to conventional methods, the proposed framework reduces average task completion time by 40% and subjective workload by 37%, while substantially improving perceived intuitiveness, user autonomy, and confidence in system performance.
本文提出一个多代理框架,使大型语言模型能够通过科学模拟模型进行受控实验,以优化制药过程设计,提高输出的具体性和实用性。
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
This work addresses the nondeterminism arising in human-in-the-loop cyber-physical systems due to human behavior, uncertainties in AI agents, and dynamic environments. It introduces, for the first time, the Reactor Model of Computation (Reactor MoC) into this domain, leveraging the Lingua Franca framework to construct a deterministic system architecture that integrates large language model–driven AI agents. Using an “intelligent driving coach” as a validation case study, the approach identifies and mitigates key challenges undermining system determinism, thereby significantly enhancing controllability and robustness. The proposed methodology offers a viable pathway toward restoring determinism in human-in-the-loop systems powered by AI agents.
本文提出使用代理仿真预测NASA TLX分数的Synthetic TLX方法,以预估技术介导任务中的人类工作负荷,并通过实验验证其有效性。