Score
Designs and runs controlled evaluations and comparison studies of different AI interaction modes (e.g., chat vs in‑code, tutor vs collaborator vs solver), including experimental protocols such as within‑subject counterbalancing and task assignments. Builds evaluation frameworks and analyzes behavioral and performance differences and short‑term carryover effects across interaction types using appropriate metrics and statistical tests.
Current AI evaluation predominantly relies on static responses, which inadequately capture the systematic capabilities of large language models in dynamic scenarios involving tool use, environmental interaction, and multi-agent collaboration. This work proposes “interactive evaluation” as a distinct paradigm, introducing a two-axis taxonomy to clarify its design principles and reporting standards. By leveraging trajectory modeling and a multi-dimensional scoring mechanism, the framework extracts evidence from interaction processes to holistically assess model performance across dimensions such as procedural fidelity, recoverability, coordination, robustness, and system-level efficacy. The study delineates core challenges in interactive evaluation and redefines the logical mapping from empirical evidence to performance judgments, thereby establishing a theoretical foundation for a unified, comparable, and interpretable next-generation AI evaluation framework.
This study addresses the challenge of effectively evaluating the impact of AI systems in knowledge work, which is hindered by traditional experimental methods that rely on unstructured textual descriptions lacking comparability, reusability, and auditability. To overcome this limitation, the authors propose the SEED framework, which formalizes human–AI collaborative experimental designs as typed participant–process graphs. This approach enables explicit representation of interaction structures, assessment of design novelty, and generation of feasible configurations under specified constraints. Integrating structured encoding, graph-guided generation, and lightweight validation, SEED significantly enhances process clarity, hypothesis specificity, and regulatory compliance in a medical triage task. The results demonstrate its effectiveness as a traceable, comparable, and generative tool for supporting rigorous experimental design in human–AI collaboration.
Existing AI evaluations predominantly rely on static benchmarking, failing to detect emergent harms—such as emotional dependency, social manipulation, and cognitive overload—that arise during prolonged human-AI interaction. To address this gap, we propose a novel dynamic evaluation paradigm centered on *interactional ethics*, moving beyond conventional single-turn output assessment. We introduce the first integrated framework for interactional ethical evaluation, combining controlled human-AI interaction experiments, NLP-based behavioral analysis, validated social science psychometric scales, and computational modeling of human impact. Our framework explicitly incorporates the evolution of human-AI relationships, downstream societal effects, and cognitive consequences as core evaluation dimensions. The project yields actionable principles for identifying interactional harms and scenario-specific evaluation protocols, thereby bridging a critical gap in safety assessment for generative AI deployed in authentic, long-term interaction settings. This work establishes a methodological foundation for context-aware, usage-oriented AI governance.
This study investigates how different interaction paradigms with large language models (LLMs) affect high school students’ performance on introductory programming tasks. A controlled experiment was conducted using ChatGPT-4o to compare three interaction modes: passive (unidirectional model output), active (model-initiated questioning), and collaborative (bidirectional human–AI negotiation). Results demonstrate that the collaborative mode significantly reduces task completion time by 37% on average, increases user satisfaction by 42%, enhances perceived helpfulness by 51%, and lowers error rates by 26%. This work provides the first empirical validation of “negotiative prompting” in programming education, establishing that interaction design—not merely model capability—is critical to learning efficacy. The findings yield a reproducible, evidence-based interaction paradigm for AI-enhanced programming instruction, offering concrete implications for pedagogical integration of LLMs in computing education.
This study addresses the challenge in current AI creativity research of simulating authentic collaborative dynamics while maintaining experimental control, due to the absence of comparable co-creative contexts. The authors introduce a controlled, turn-taking alternate uses test platform that, for the first time, enables direct comparison between human–human and human–AI interactive co-creativity under rigorously matched conditions, incorporating a “creative seed” intervention mechanism. Integrating a GPT-4 interaction system, a behavioral experimentation platform, psychological scales (e.g., BAS Drive), and multidimensional creativity metrics, the research reveals that GPT-4 partners generate ideas of originality comparable to human partners within identical time constraints. Furthermore, individual motivation moderates the positive effect of interaction on originality, cognitive offloading diminishes originality in human partners, and prior exposure to highly creative ideas significantly enhances subsequent performance. This framework disentangles three classes of influences—participant traits, partner perception, and content dynamics—offering a novel paradigm for AI co-creativity research.
This study addresses a critical gap in the evaluation of AI tutoring systems, which has traditionally emphasized the instructional quality of feedback while overlooking how students actually engage with and utilize it. The authors propose a novel assessment framework that integrates student behavioral data, introducing interaction-based dimensions to examine whether learners adopt and correctly apply AI-generated feedback. Through large-scale analysis of code submissions and interaction logs, the research demonstrates that behavioral signals are more effective than conventional instructional quality metrics in predicting students’ perceived usefulness of feedback. The framework’s validity is further confirmed across two consecutive semesters in authentic classroom settings, offering a new paradigm for the comprehensive evaluation of AI tutoring systems.
This study addresses the lack of infrastructure supporting reproducible, longitudinal, and real-time human–AI collaboration experiments, which has hindered systematic investigation into how design attributes of AI teammates influence team trust, coordination, and decision-making. To bridge this gap, the authors introduce TRAIL, a novel platform that embeds configurable and reproducible AI teammates within an instrumented, authentic collaborative environment, enabling longitudinal experimentation and behavioral analysis. TRAIL innovatively integrates the Big Five personality model, selective messaging channels, a dual-memory architecture, chained experimental scheduling, and textual similarity analysis tools to systematically modulate AI personality, communication timing, and interaction style. In a six-round classroom study with 51 students, TRAIL sustained stable AI collaboration and revealed significant differential effects of AI personality on team perceptions of contribution, linguistic alignment, group atmosphere, and reliance on the AI teammate.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.