Score
Designs, builds, and analyzes sociotechnical and ML-enabled workflows that explicitly anticipate and manipulate how predictions, prompts, or policies change human behavior by creating decision rules, experimental treatments, runtime fallback mechanisms, and task-allocation schemes between AI and workers. Work includes specifying controlled and edge intervention experiments, simulating and testing interventions, implementing deployment and training interventions and engagement policies, and performing causal and trade-off analyses (for example, current output versus future skill investment) to evaluate and refine intervention strategies.
This study addresses the conceptual ambiguity surrounding terms such as “human-in-the-loop” (HITL) and “human-on-the-loop” (HOTL) in contemporary AI systems, which impedes interdisciplinary collaboration and generates regulatory uncertainty. The authors propose a causal-structure-based classification framework that clearly distinguishes between constitutive human involvement (HITL) and corrective oversight (HOTL), further delineating three temporal modes of HOTL and their associated dimensions of cognitive integration. Through conceptual analysis, causal modeling, and cross-disciplinary theoretical synthesis, the work develops a typology of human participation that separates descriptive from normative considerations. This framework illuminates the systemic design challenges arising from role duality and underscores “effective intervention capacity” as the cornerstone of meaningful human oversight, thereby offering a theoretical foundation for the coordinated governance of legal, technical, and ethical dimensions in AI.
This study addresses the critical gap in effective tools for predicting employees’ psychological and behavioral responses to AI integration in knowledge work, which hinders precise management of workforce transformation. It proposes a novel computational framework that, for the first time, integrates generative agents with organizational behavior theory to construct dynamic employee agents powered by large language models. These agents synthesize human resource records, psychometric assessments, and digital behavioral logs to simulate individual cognitive, emotional, and behavioral trajectories during organizational change, all within a multi-layered privacy-preserving architecture. The resulting infrastructure offers a scalable and ethically grounded platform for forward-looking simulations, thereby enabling responsible, AI-informed workforce policy design.
This study addresses the challenge of effectively evaluating the impact of AI systems in knowledge work, which is hindered by traditional experimental methods that rely on unstructured textual descriptions lacking comparability, reusability, and auditability. To overcome this limitation, the authors propose the SEED framework, which formalizes human–AI collaborative experimental designs as typed participant–process graphs. This approach enables explicit representation of interaction structures, assessment of design novelty, and generation of feasible configurations under specified constraints. Integrating structured encoding, graph-guided generation, and lightweight validation, SEED significantly enhances process clarity, hypothesis specificity, and regulatory compliance in a medical triage task. The results demonstrate its effectiveness as a traceable, comparable, and generative tool for supporting rigorous experimental design in human–AI collaboration.
This study investigates the cognitive mechanisms underlying teachers’ design of multi-agent instructional workflows, reconceptualizing AI-TPACK (Artificial Intelligence–Technological Pedagogical Content Knowledge) as a dynamic cognitive practice rather than a static knowledge construct. Drawing on a mixed-methods approach that integrates cluster analysis and Markov chain modeling, the research analyzes behavioral logs from 61 teachers, 15 instructional artifacts, and 12 interviews to identify three distinct design archetypes: Systematic Optimizers, Prolific Creators, and Passive Observers. The findings elucidate how systems thinking, pedagogical beliefs, and self-efficacy dynamically shape the integration of AI-TPACK, offering empirical grounding for differentiated support strategies in professional development for intelligent educational technologies.
This study addresses the frequent failures or delays enterprises encounter when deploying machine learning models, which often stem from non-technical factors such as organizational and managerial challenges. Drawing on 66 hours of practitioner talks from the MLOps community, the research employs qualitative methods—specifically manual coding and thematic analysis—to systematically identify 17 socio-technical anti-patterns that impede the successful operationalization of ML systems. Moving beyond a purely technical lens, the work emphasizes team collaboration and organizational mechanisms, offering targeted recommendations spanning technology, processes, and structural arrangements. These insights are further validated through cross-referencing with existing literature, thereby providing a systematic and actionable framework to guide ML engineering practices in real-world settings.
This study investigates how human intervention influences service outcomes when autonomous AI fails in customer service contexts, with a focus on the cognitive and emotional dimensions of intervention efficacy. Drawing on a randomized field experiment conducted on Alibaba’s Taobao platform, the research compares customer service agents’ performance in chat tasks with and without AI assistance. It provides the first empirical evidence that intervention effectiveness depends on the type of AI failure (technical versus emotional), the timing of intervention, and the level of subsequent human effort. Findings reveal that while AI reduces average chat duration, it lowers satisfaction ratings for conversations it handles. Human intervention effectively preserves service quality in technical failures but shows limited efficacy in emotional failures. Early intervention not only enhances agents’ subsequent effort but also significantly improves service outcomes for AI-unmanageable chats, generating positive spillover effects through workflow adaptation.
Current AI alignment approaches rely on static human preferences, which struggle to capture the dynamic, context-dependent nature of human–AI collaboration. This work proposes a paradigm shift toward “interactive complementarity,” emphasizing that preferences emerge dynamically through the co-evolution of human and AI behaviors. Introducing a trajectory-level perspective on dynamic alignment, the study integrates machine learning with social science theories and insights from interdisciplinary workshops to construct a framework for modeling human–AI interaction dynamics. This framework reveals novel forms of asymmetry and coordination challenges inherent in such systems. By establishing an alignment agenda tailored to dynamic human–AI workflows, the research lays a theoretical foundation for developing AI systems capable of interactive adaptation.
This study addresses the fragmented understanding of sociotechnical risks in human-AI collaboration, which has hindered the identification of common failure mechanisms and effective interventions. To overcome this limitation, the work proposes a unified lifecycle framework encompassing four phases: task allocation, interaction, feedback, and adoption. Through a cross-domain literature review and conceptual modeling, it synthesizes empirical evidence from healthcare, journalism, education, and scientific research to establish the first comprehensive risk taxonomy. The analysis reveals six core risk clusters—including miscalibrated trust, cognitive overload, and responsibility gaps—and elucidates their cascading interdependencies. By moving beyond isolated risk assessments, this research provides a theoretical foundation for resilient human-AI collaboration and informs end-to-end governance strategies and human-centered AI system design.