Score
Designs and evaluates systems that detect low-confidence model outputs and implement abstention or human-in-the-loop review, including mechanisms to route uncertain samples to human reviewers, set confidence thresholds, and prioritize examples for review. Analyzes trade-offs between reduced misclassification and human annotation cost and builds workflows and metrics to minimize human effort while meeting target reliability.
This work addresses the trade-off between human annotation cost and system accuracy in human-in-the-loop classification by proposing an optimization framework based on a dual-threshold strategy. By setting upper and lower confidence thresholds, the system automatically processes high-certainty samples while routing only ambiguous cases to human reviewers. The approach formalizes the human-AI collaboration problem, identifies the critical region where human intervention yields diminishing returns, and quantifies the marginal benefit of manual review across diverse scenarios through probabilistic score modeling, Monte Carlo simulation, and optimization algorithms. Empirical evaluations demonstrate the framework’s generality and effectiveness across multiple domains—including entity resolution, fraud detection, medical triage, and content moderation—achieving high accuracy while substantially reducing human workload.
This study addresses the challenging human supervision task of verifying the factual accuracy of AI-generated outputs. Methodologically, we propose a human-AI collaborative fact-checking framework featuring an AI confidence estimation and explanation generation module; it dynamically integrates AI scores with human judgments while modulating human trust by presenting only verifiable search evidence—not definitive conclusions. Our key contribution is the first systematic empirical validation that “lightweight AI assistance”—defined as providing only auditable evidence—significantly mitigates human overreliance on AI. Results demonstrate that our fusion mechanism improves human verification accuracy by 12.7% over both pure-human and pure-AI baselines. These findings establish a scalable technical pathway and cognitively grounded design principles for building trustworthy AI supervision paradigms.
Current human baselines in large language model evaluation lack methodological rigor and transparency, undermining the validity of claims such as “superhuman performance.” Method: This paper pioneers the systematic integration of classical measurement theory into AI evaluation, establishing a comprehensive theoretical framework spanning human baseline design, execution, and reporting. It introduces an actionable quality assessment system and a standardized checklist, derived via meta-review–driven framework development, structured checklist design, and empirical systematic auditing. Contribution/Results: Applying this framework to diagnose 115 human baseline studies, we identify pervasive methodological flaws. The resulting open-source audit tool significantly enhances reproducibility, comparability, and accountability in AI evaluation. By grounding benchmarking practice in psychometric principles, our work provides a rigorous methodological foundation for scientifically credible model capability assessment.
Systematic Literature Review (SLR) updates face a critical trade-off between reducing human effort in study screening and preserving evidence completeness. Method: We propose and empirically evaluate a human-in-the-loop screening framework for SLR updates in software engineering, employing Random Forest and SVM models optimized for 100% recall during pre-screening. Contribution/Results: Our approach reduces manual screening effort by 33.9% while achieving an F1-score of 0.33—confirming its unsuitability as a fully automated replacement but strong value as a high-recall pre-screening tool. Dual-reviewer initial screening yielded results closest to final inclusion decisions. This work presents the first rigorous, recall-guaranteed application of machine learning in SLR updating, coupled with quantitative evaluation of human-AI collaboration efficiency. It provides a reproducible, generalizable methodological foundation for evidence-driven automation of SLRs.
This study addresses the stagnation of enterprise AI initiatives in regulated financial institutions due to the absence of quantifiable evaluation criteria. Focusing on six document-intensive workflows, it systematically compares AI system performance across four model families and three tool configurations, distinguishing between demonstration and production environments. For the first time, it links deployment feasibility with human review rates. The authors propose a production-grade evaluation framework encompassing accuracy, reproducibility, traceability, and informative confidence, integrating multi-model comparison, confidence signals, source citation, and self-verification mechanisms. Experiments reveal that 56.1% of the 72 evaluated configurations meet production readiness thresholds. Incorporating source citation and confidence estimation reduces human review requirements to 49%, and adding self-verification further lowers this to 44%, albeit at the cost of reduced error tolerance.
Traditional systematic literature reviews suffer from low efficiency and susceptibility to selection bias during clinical trial screening and data extraction. This work proposes two task-oriented multi-agent systems (MAS)—one for automated trial screening and another for structured information extraction—both integrating human-in-the-loop mechanisms to support clinical decision-making. The systems innovatively employ heterogeneous large language model agents, multi-round cross-review, standardized workflows, retrieval-augmented context control, and iterative error correction, substantially enhancing accuracy and scalability. In a real-world network meta-analysis replication, the approach not only fully reproduced all trials included in the original study but also identified additional eligible trials missed by manual screening, thereby updating the clinical conclusions.
Current mechanisms struggle to verify whether revisions to scientific manuscripts substantively address peer reviewers’ concerns with supporting textual evidence. This work proposes AutoSupervision, a novel framework that leverages transparent peer review records from 56,000 papers in Nature Communications to construct a closed-loop evaluation system powered by large language models (e.g., GPT-5.5). The system automatically identifies reviewer concerns, assesses the effectiveness of author revisions, and locates supporting evidence within the revised text. Experimental results show that large language models achieve strong performance in identifying reviewer concerns (F1 = 0.754), yet evidence-based verification remains challenging, with the best current model attaining only an F1 score of 0.501. This study establishes a verifiable, evidence-driven paradigm for automated assessment of scientific manuscript revisions.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.