automated response scoring

Design and implement systems that automatically apply per-question evaluation rubrics to model or user-generated responses and compute per-response quality and safety scores; include automated detection/flagging of low‑quality or harmful answers and workflows to scale these evaluations across many models or large volumes of responses.

automatedresponsescoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$229K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of evaluating automatically generated multiple-choice questions, specifically the difficulty of quality assessment, inaccurate defect detection, and uncertain revision effectiveness. Employing a narrative review and NLP benchmark auditing, this work proposes a quality assurance framework conceptualized as a sequence of independently verified decisions. The research constructs a 19-criterion mapping system that distinguishes surface-level inspection from content-level judgment, advocating for quality assurance to be treated as an independent verification process. Furthermore, it reveals a paradox wherein high label accuracy coexists with weak positive-instance detection, noting that current evidence remains insufficient to establish definitive repair effects. Ultimately, this paper offers a novel evaluation paradigm and practical guidelines for automated question generation.

automated assessmentitem-writing flaw detectionmultiple-choice questions

This work addresses the lack of structured mechanisms capable of dynamically evaluating and guiding behavior as large language models evolve toward open-ended autonomous agents. It proposes rubrics as a unified framework to translate complex quality judgments into structured, actionable specifications, systematically elucidating their progressive roles across evaluation, training, and internal agent behavior for the first time. By designing structured rubrics, decomposing assessments into multiple dimensions, generating dense feedback, and analyzing self-improvement behaviors, the study demonstrates the reliability of rubrics in ensuring generation quality, execution fidelity, adherence to theoretical constraints, and mitigation of safety threats. Furthermore, it establishes a cross-domain benchmarking framework that bridges human intent with machine behavior.

Autonomous AgentsEvaluation FrameworkHuman-AI Alignment

This study addresses the challenge that developer-provided usefulness labels for LLM-generated code review comments in industrial settings are often biased by workflow pressures and organizational factors, undermining their reliability as ground truth. Leveraging 2,604 LLM-generated comments and corresponding engineer annotations from Beko, this work presents the first systematic evaluation in a real-world industrial context of the alignment between human feedback and two automated evaluation paradigms: G-Eval and LLM-as-a-Judge. The experiments span multiple models—including Gemini-2.5-pro, GPT-4.1-mini, and GPT-5.2—and incorporate qualitative validation through interviews with engineering leads. Results reveal only moderate agreement (0.44–0.62) between automated metrics and human labels, with performance sensitive to both model choice and evaluation design, thereby exposing the limitations of relying solely on automated assessments and challenging the common assumption that developer annotations constitute a gold standard.

automated code reviewdeveloper feedbackevaluation reliability

This work addresses the lack of a unified framework in existing LLM evaluation methods based on rubrics, which suffer from terminological inconsistency and fragmented implementations. We propose Autorubric, an open-source framework that systematically integrates analytic rubrics, single- and multi-rater aggregation, few-shot calibration, bias mitigation, and psychometric reliability metrics such as Cohen’s κ. Autorubric supports binary, ordinal, and nominal criteria and provides actionable default configurations. Experimental results demonstrate that Autorubric achieves 80% and 87% accuracy on RiceChem and CHARM-100, respectively; successfully improves peer-review agent scores from 0.47 to 0.85; and significantly enhances AdvancedIF performance via RL-based rewards (+0.039, p=0.032).

bias mitigationfew-shot calibrationLLM evaluation

Latest Papers

What's happening recently
View more

This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.

automated researchevaluation benchmarksresearch evaluation

This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.

AI evaluationcontextual alignmenthuman-aligned scoring

This work addresses the fragility of existing safety evaluation models when confronted with variations in scoring criteria and prompts, which undermines their ability to consistently adhere to diverse judgment standards. The authors frame safety assessment as a criterion-following problem and propose a curriculum learning framework that progresses from “reliable” to “expressive” behaviors. By integrating dynamically generated instance-conditional scoring rubrics with supervised fine-tuning, they train a 12B-parameter language model to robustly align with shifting evaluation guidelines. Their approach is the first to systematically resolve judgment instability under varying criteria, achieving accuracies of 94.12%–94.88% across three distinct scoring standards with a remarkably low cross-criterion performance variance of only 0.76—significantly outperforming general-purpose large language models, specialized safety classifiers, and reasoning-based evaluators with fewer than 30B parameters.

evaluation robustnessfalse negative rateprompt variation

Existing AI safety evaluation benchmarks suffer from redundancy, high inter-correlation, and potential sandbagging—where models deliberately underperform—rendering aggregated scores difficult to interpret and trust. This work presents the first large-scale application of Item Response Theory (IRT) to safety assessment of large language models (LLMs), integrating factor analysis and adaptive testing to distill three interpretable latent traits from multiple benchmarks: refusal strictness, truthfulness, and contextual harm. The proposed IRT-based framework achieves 97–99% fidelity in replicating individual benchmark outcomes using only about ten adaptively selected items, substantially improving evaluation efficiency. Furthermore, it effectively detects sandbagging behaviors, demonstrating IRT’s reliability, scalability, and auditability in LLM safety evaluation.

AI safetyevaluation reliabilitylanguage models

Hot Scholars

MB

Michael Backes

Chairman and Founding Director of the CISPA Helmholtz Center for Information Security
SecurityprivacycryptographyAI
UN

Usman Naseem

Lecturer (Asst. Prof.) @Macquarie University
Natural Language ProcessingLLM AlignmentNLP for Social GoodTrust and Safety
DZ

Ding Zhao

Carnegie Mellon University
Trustworthy AIAI safetyreinforcement learningautonomous vehicles
MN

Mehwish Nasim

Senior Lecturer UWA Australia, Network Analysis & Social Influence Modelling (NASIM)Lab
Information WarfareNetwork ScienceNLPComplex Systems
PN

Ping Nie

Waterloo University
Natural Language ProcessingInformation RetrievalRecommendation SystemsTime Series Forecasting