design human-in-the-loop systems

Designs and implements end-to-end systems that integrate human reviewers with automated models to guide generation, curate and synthesize training data, perform validation and verification, and evaluate model outputs. This work specifies intervention and adjudication protocols, routes ambiguous cases and human review checkpoints through automated workflows, builds human-guided data-synthesis and refinement loops, composes hybrid human–LLM scoring or LLM-as-judge pipelines, and defines aggregation and workflow strategies to minimize human burden while ensuring reliable evaluation and learning.

designhuman-in-the-loopsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

EvalAssist: A Human-Centered Tool for LLM-as-a-Judge

Jul 02, 2025
ZA
Zahra Ashktorab
🏛️ IBM Research

To address the high cost, inconsistent criteria, and low efficiency of evaluating LLM outputs—particularly in multi-model and multi-prompt selection—this paper proposes a human-centered automated evaluation framework. Methodologically: (1) we design an interactive, web-based environment for developing structured, shareable evaluation specifications; (2) we integrate prompt chaining with a dedicated harmful-content detection model to implement a lightweight, portable LLM-as-a-judge pipeline within the open-source UNITXT library; (3) the framework requires no fine-tuning, relying solely on off-the-shelf LLMs and custom prompt engineering. Deployed internally, it serves hundreds of users, substantially reducing human evaluation effort and turnaround time. Empirical adoption demonstrates improved consistency, reproducibility, and standardization across diverse models and tasks.

Detecting harms and risks in LLM outputs effectivelyReducing time and cost in LLM-as-a-judge workflowsSimplifying evaluation of diverse LLM outputs for tasks

Maximizing Signal in Human-Model Preference Alignment

Mar 06, 2025
KK
Kelsey Kraus
🏛️ Cisco Systems

To address the challenges of aligning large language model (LLM) outputs with end-user preferences and mitigating high noise in human feedback, this paper proposes a Noise–Signal Decoupling Framework that systematically disentangles stochastic annotation noise from genuine preference signals within labeling disagreements. Methodologically, it introduces a human-feedback-based annotation quality analysis and consistency modeling mechanism, designs a preference-driven supervised fine-tuning strategy, and incorporates dual guardrail classifiers for closed-loop evaluation. Its key innovation lies in explicitly formulating preference signal maximization as an optimization objective—integrated throughout both training and evaluation. Experiments demonstrate significant noise reduction: the approach improves accuracy and fairness on user-consensus–sensitive tasks—including toxicity detection and key-point extraction in summarization—while yielding a reusable, trustworthy evaluation practice guideline.

Challenges in evaluating LLM outputs due to creativity and fluency.Maximizing signal in human feedback for user-aligned model behavior.Need for human judgments in model training and evaluation.

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Jun 26, 2024
AB
Anna Bavaresco
🏛️ University of Amsterdam | University of Trento | University of Copenhagen | Utrecht University | ETH Zürich | Saarland University | Universidade de Lisboa | LMU Munich | University of Potsdam | Heriot-Watt University | Unbabel | MCML

This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.

Assessing validity of LLMs replacing human judges in NLP evaluationsEvaluating reproducibility of proprietary vs open-weight LLM modelsMeasuring variance in LLM performance across diverse NLP tasks

This study investigates the mechanisms influencing human–LLM judgment alignment in human–AI collaborative evaluation, focusing on how task characteristics and AI assistance strategies shape users’ construction and dynamic refinement of evaluation criteria, as well as their model selection behavior. Method: We conducted a controlled human–AI interaction study involving 15 ML practitioners performing 131 real-world evaluation tasks, comparing direct assessment versus pairwise comparison paradigms, augmented by multi-round LLM-assisted judgments and qualitative behavioral analysis. Contribution/Results: We present the first empirical evidence that direct assessment significantly enhances user engagement and criterion-task alignment, facilitating personalized criterion customization, dynamic judgment adjustment, and adaptive model switching. Based on these findings, we propose design principles for front-end evaluation tools tailored to human–AI collaboration. Our work advances low-overhead, interpretable, and task-adaptive AI-assisted evaluation frameworks.

Aligning human and LLM judgments for evaluationsImproving AI-assisted evaluation strategies and toolsReducing cost and time in LLM output assessments

Latest Papers

What's happening recently
View more

Current mechanisms struggle to verify whether revisions to scientific manuscripts substantively address peer reviewers’ concerns with supporting textual evidence. This work proposes AutoSupervision, a novel framework that leverages transparent peer review records from 56,000 papers in Nature Communications to construct a closed-loop evaluation system powered by large language models (e.g., GPT-5.5). The system automatically identifies reviewer concerns, assesses the effectiveness of author revisions, and locates supporting evidence within the revised text. Experimental results show that large language models achieve strong performance in identifying reviewer concerns (F1 = 0.754), yet evidence-based verification remains challenging, with the best current model attaining only an F1 score of 0.501. This study establishes a verifiable, evidence-driven paradigm for automated assessment of scientific manuscript revisions.

grounded evidencepeer reviewreviewer feedback

Traditional peer review faces scalability bottlenecks, while large language model (LLM)-driven automated review lacks systematic investigation into its reliability, robustness, and security. This work addresses this gap by offering the first system-oriented analysis, focusing on two core tasks: critique generation and score prediction. It establishes a taxonomy of LLM-based reviewing approaches and comprehensively evaluates key technical strategies, including prompt engineering, supervised fine-tuning, retrieval augmentation, and alignment optimization. The study uncovers emerging security threats such as prompt injection and data poisoning, examines challenges arising from subjective disagreement and cross-domain generalization, and highlights limitations and domain biases in current benchmarks. Building on these insights, the paper proposes a roadmap toward developing reliable, transparent, and trustworthy AI-assisted scientific review systems.

LLM-based peer reviewreliabilityrobustness

Hot Scholars

YL

Yunyao Li

Director of Machine Learning, Adobe Experience Platform
Natural Language ProcessingMachine LearningHuman Computer InteractionData Management
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
MV

Mihaela van der Schaar

University of Cambridge, The Alan Turing Institute
machine learningML for healthcarecompression and streamingmulti-user networking
PL

Philippe Laban

Senior Research Scientist, Microsoft Research
nlphcifactualityllm evaluation
AK

Alois Knoll

Technische Universität München
RoboticsAISensor Data FusionAutonomous Driving