Score
Designs, implements, and administers human-subject studies and perceptual experiments to collect subjective ratings, preferences, and expert judgments about system outputs, including recruitment, rating interfaces, and blind or comparative protocols. Builds evaluation protocols and guidelines (e.g., preference tests, rater instructions), and analyzes the resulting data to measure agreement, reliability, statistical significance, preference aggregation, and to derive human-grounded evaluation metrics such as perceptual quality or bias assessments.
This study investigates whether large language model (LLM)-based agents can reliably emulate human subjective ratings of data visualization designs. Method: Through a three-stage empirical study, we systematically assess agent–human rating alignment, identify influencing factors, and evaluate enhancement strategies. We propose a multimodal LLM agent framework integrating visualization preprocessing, structured prompt engineering, and domain-knowledge injection. Contribution/Results: We uncover that expert prior confidence—rather than visualization features or prompt design—is the strongest predictor of human–agent alignment. Contrary to common practice, standard prompt-enhancement techniques introduce systematic biases. Under high-expert-confidence assumptions, the agent enables rapid prototype evaluation but must serve as a complement—not a substitute—for human-centered user studies. To our knowledge, this is the first work to systematically characterize alignment patterns and boundary conditions of LLMs in subjective visualization assessment.
This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
This study investigates how human-in-the-loop (HITL) feedback influences users’ perceptions of system accuracy and trust, highlighting the critical moderating role of task subjectivity. Through three controlled user experiments that systematically differentiate between objective and subjective task contexts, the research analyzes behavioral measures to assess the effects of feedback interaction. Findings reveal that in objective tasks, providing feedback significantly diminishes users’ trust in and perceived accuracy of the system, whereas this negative effect vanishes in subjective tasks. These results underscore task type as a pivotal factor shaping human–AI trust dynamics and offer important theoretical grounding and practical guidance for the design of HITL systems.
This study investigates the mechanisms influencing human–LLM judgment alignment in human–AI collaborative evaluation, focusing on how task characteristics and AI assistance strategies shape users’ construction and dynamic refinement of evaluation criteria, as well as their model selection behavior. Method: We conducted a controlled human–AI interaction study involving 15 ML practitioners performing 131 real-world evaluation tasks, comparing direct assessment versus pairwise comparison paradigms, augmented by multi-round LLM-assisted judgments and qualitative behavioral analysis. Contribution/Results: We present the first empirical evidence that direct assessment significantly enhances user engagement and criterion-task alignment, facilitating personalized criterion customization, dynamic judgment adjustment, and adaptive model switching. Based on these findings, we propose design principles for front-end evaluation tools tailored to human–AI collaboration. Our work advances low-overhead, interpretable, and task-adaptive AI-assisted evaluation frameworks.
This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.
This work addresses the high cost, lengthy timelines, and limited fidelity of traditional UI/UX evaluation methods—such as user studies and A/B testing—in simulating authentic user feedback. To overcome these limitations, the authors propose PerceptUI, a novel framework that enables fine-grained, persona-based UI/UX feedback generation for the first time. Leveraging multimodal large language models, PerceptUI integrates persona-conditioned prompting, contrastive reflection-based fine-tuning, and a failure-trajectory-driven prompt evolution mechanism to distill rational justifications from human decision-making processes, thereby enhancing the model’s introspective capabilities. Experimental results demonstrate that PerceptUI achieves human-level authenticity in generated feedback across multiple domains and datasets, generalizes effectively to unseen interface issues and user personas, and supports the synthesis of population-level response distributions.
This work addresses a key limitation of current large language models (LLMs) as AI evaluators—their reliance on aggregate consensus while neglecting individual judgment variability. The study presents the first systematic exploration of simulating personalized human preferences by integrating evaluator-specific auxiliary data, such as chain-of-thought reasoning traces and interface telemetry, with in-context learning. Findings reveal that a neutral usage tendency emerges as a stable and predictable cross-task indicator of individual preference. The proposed approach achieves up to a 9.9-percentage-point improvement over baseline methods, with reasoning traces contributing the largest performance gain, whereas interface telemetry often degrades accuracy. These results highlight both the promise and inherent constraints of personalizing LLM-based evaluation.
This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.
This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.