🤖 AI Summary
This study addresses the challenges of evaluating automatically generated multiple-choice questions, specifically the difficulty of quality assessment, inaccurate defect detection, and uncertain revision effectiveness. Employing a narrative review and NLP benchmark auditing, this work proposes a quality assurance framework conceptualized as a sequence of independently verified decisions. The research constructs a 19-criterion mapping system that distinguishes surface-level inspection from content-level judgment, advocating for quality assurance to be treated as an independent verification process. Furthermore, it reveals a paradox wherein high label accuracy coexists with weak positive-instance detection, noting that current evidence remains insufficient to establish definitive repair effects. Ultimately, this paper offers a novel evaluation paradigm and practical guidelines for automated question generation.
📝 Abstract
Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.