Score
Designs and implements the systems, processes, and tooling for recruiting, onboarding, training, scheduling, compensating, monitoring, and retaining human raters or annotators, as well as workflows for task assignment, quality control, and data security. Builds analytics and controls to measure rater quality, throughput, disagreement, bias, and capacity, and uses those analyses to adjust training, calibration, and pool composition.
This study addresses the opacity and accountability challenges in AI-powered hiring systems, which stem from their complex supply chains that obscure the origins of algorithmic bias. Through regulatory analysis, system dependency modeling, and a multi-stakeholder perspective—complemented by case studies and an examination of implementation ambiguities—the work demonstrates for the first time that bias arises primarily from interactions among system components rather than from isolated modules. It further identifies a structural contradiction: deploying organizations bear legal responsibility yet lack technical visibility into upstream components. The research pinpoints two core barriers to effective bias assessment and accountability and proposes a holistic, supply-chain-wide governance framework featuring system-level audits, vendor guidelines, continuous monitoring, and cross-component documentation.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
Early-stage candidate screening suffers from low efficiency due to the need to integrate heterogeneous information—resumes, interview videos, coding assignments, and public online data. This paper proposes a risk-aware, modular multi-agent system orchestrated by constraint-augmented large language models (LLMs), encompassing multimodal parsing (PDF/video), structured profile construction, knowledge-graph–driven public evidence verification, dual-dimension (technical/cultural) scoring with explicit risk penalization, and a human-in-the-loop review interface. Key contributions include: (i) the first explainable and traceable risk-aware scoring framework; (ii) a novel efficiency metric—“time per qualified candidate”; and (iii) component-level attribution and low-variance decision-making. Evaluated on real screening of 64 Python backend engineers, the system reduced time per qualified candidate from 3.33 to 1.70 hours, maintaining baseline precision and recall, while preserving final hiring authority exclusively with human recruiters.
This paper addresses systemic unfairness in AI-driven recruitment—manifesting as ranking bias and inaccurate interview evaluations—stemming from the propagation of human biases. It proposes the first end-to-end fairness analysis framework for recruitment AI. Methodologically, it explicitly disentangles bias sources across data, algorithm, and deployment layers, integrating a socio-technical systems perspective with statistical fairness metrics (Demographic Parity, Equalized Odds), three categories of bias mitigation strategies, and a third-party audit toolchain. Key contributions include: (1) taxonomizing 12 canonical bias scenarios; (2) constructing a fairness evaluation matrix comprising 27 operational metrics; and (3) proposing organization-aware, co-optimization pathways balancing fairness and operational efficacy. The work establishes a theoretical analytical paradigm for academia and delivers an actionable governance roadmap for industry, bridging critical gaps in cross-layer bias attribution and real-world impact assessment.
This study addresses bias in skill-oriented job matching systems that may undermine hiring fairness. The authors propose a unified two-stage governance framework: in the first stage, a chatbot extracts candidate skills and disentangles hard and soft constraint biases; in the second stage, preferences from candidates, employers, and regulators are integrated through a multi-stakeholder recommendation mechanism grounded in social choice theory. The framework incorporates distributional auditing, counterfactual testing, and dynamic fairness evaluation to enable auditable bias detection. It automatically triggers corrective actions or generates compliance reports when predefined fairness thresholds are violated, thereby significantly enhancing the system’s fairness, transparency, and regulatory alignment—such as with the EU AI Act.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
This study addresses the challenge of low-quality bug reports in crowdsourced testing, which impose substantial review burdens on developers and lack effective mechanisms to improve tester performance. The authors propose a large language model–based multi-agent evaluation framework that automatically assesses reports along three dimensions—textuality, sufficiency, and competitiveness—and integrates actionable feedback into human workflows. Through a four-phase controlled experiment combined with mixed-methods analysis, they provide the first empirical evidence that evaluative agents not only serve as post-hoc adjudicators but also function as in-process feedback sources, significantly enhancing the quality of report revisions, improving first-submission performance in subsequent tasks, and facilitating cross-application knowledge transfer. User studies further confirm the intelligibility and practical utility of the generated feedback.