Score
Designs and implements end-to-end systems that integrate human reviewers with automated models to guide generation, curate and synthesize training data, perform validation and verification, and evaluate model outputs. This work specifies intervention and adjudication protocols, routes ambiguous cases and human review checkpoints through automated workflows, builds human-guided data-synthesis and refinement loops, composes hybrid human–LLM scoring or LLM-as-judge pipelines, and defines aggregation and workflow strategies to minimize human burden while ensuring reliable evaluation and learning.
This work addresses persistent cooperation barriers and evolutionary pathways in human–large language model (LLM) collaboration. Methodologically, it integrates game-theoretic modeling, cognitive science principles, multi-agent frameworks, and empirical behavioral analysis of LLMs to conduct a cross-paradigm comparative study. The contributions include: (1) the first unified taxonomy of human–model collaboration spanning the entire lifecycle, with a formally defined autonomous agent–based collaboration paradigm; (2) the construction of the inaugural human–model collaboration knowledge graph; and (3) the systematic identification of six fundamental open challenges—namely, interpretability, objective alignment, dynamic adaptability, trust calibration, value consistency, and cooperative agency emergence. Collectively, these results establish a rigorous theoretical foundation and methodological framework for designing trustworthy, robust, and goal-coherent human–LLM collaborative systems.
To address the high cost, inconsistent criteria, and low efficiency of evaluating LLM outputs—particularly in multi-model and multi-prompt selection—this paper proposes a human-centered automated evaluation framework. Methodologically: (1) we design an interactive, web-based environment for developing structured, shareable evaluation specifications; (2) we integrate prompt chaining with a dedicated harmful-content detection model to implement a lightweight, portable LLM-as-a-judge pipeline within the open-source UNITXT library; (3) the framework requires no fine-tuning, relying solely on off-the-shelf LLMs and custom prompt engineering. Deployed internally, it serves hundreds of users, substantially reducing human evaluation effort and turnaround time. Empirical adoption demonstrates improved consistency, reproducibility, and standardization across diverse models and tasks.
To address the challenges of aligning large language model (LLM) outputs with end-user preferences and mitigating high noise in human feedback, this paper proposes a Noise–Signal Decoupling Framework that systematically disentangles stochastic annotation noise from genuine preference signals within labeling disagreements. Methodologically, it introduces a human-feedback-based annotation quality analysis and consistency modeling mechanism, designs a preference-driven supervised fine-tuning strategy, and incorporates dual guardrail classifiers for closed-loop evaluation. Its key innovation lies in explicitly formulating preference signal maximization as an optimization objective—integrated throughout both training and evaluation. Experiments demonstrate significant noise reduction: the approach improves accuracy and fairness on user-consensus–sensitive tasks—including toxicity detection and key-point extraction in summarization—while yielding a reusable, trustworthy evaluation practice guideline.
This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.
This study investigates the mechanisms influencing human–LLM judgment alignment in human–AI collaborative evaluation, focusing on how task characteristics and AI assistance strategies shape users’ construction and dynamic refinement of evaluation criteria, as well as their model selection behavior. Method: We conducted a controlled human–AI interaction study involving 15 ML practitioners performing 131 real-world evaluation tasks, comparing direct assessment versus pairwise comparison paradigms, augmented by multi-round LLM-assisted judgments and qualitative behavioral analysis. Contribution/Results: We present the first empirical evidence that direct assessment significantly enhances user engagement and criterion-task alignment, facilitating personalized criterion customization, dynamic judgment adjustment, and adaptive model switching. Based on these findings, we propose design principles for front-end evaluation tools tailored to human–AI collaboration. Our work advances low-overhead, interpretable, and task-adaptive AI-assisted evaluation frameworks.
Current mechanisms struggle to verify whether revisions to scientific manuscripts substantively address peer reviewers’ concerns with supporting textual evidence. This work proposes AutoSupervision, a novel framework that leverages transparent peer review records from 56,000 papers in Nature Communications to construct a closed-loop evaluation system powered by large language models (e.g., GPT-5.5). The system automatically identifies reviewer concerns, assesses the effectiveness of author revisions, and locates supporting evidence within the revised text. Experimental results show that large language models achieve strong performance in identifying reviewer concerns (F1 = 0.754), yet evidence-based verification remains challenging, with the best current model attaining only an F1 score of 0.501. This study establishes a verifiable, evidence-driven paradigm for automated assessment of scientific manuscript revisions.
Traditional peer review faces scalability bottlenecks, while large language model (LLM)-driven automated review lacks systematic investigation into its reliability, robustness, and security. This work addresses this gap by offering the first system-oriented analysis, focusing on two core tasks: critique generation and score prediction. It establishes a taxonomy of LLM-based reviewing approaches and comprehensively evaluates key technical strategies, including prompt engineering, supervised fine-tuning, retrieval augmentation, and alignment optimization. The study uncovers emerging security threats such as prompt injection and data poisoning, examines challenges arising from subjective disagreement and cross-domain generalization, and highlights limitations and domain biases in current benchmarks. Building on these insights, the paper proposes a roadmap toward developing reliable, transparent, and trustworthy AI-assisted scientific review systems.