A Benchmark for LLM's Understanding of Middle School and High School Science Topics

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of standardized evaluation tools for large language models (LLMs) in K-12 science education by constructing a middle school science benchmark aligned with the Next Generation Science Standards (NGSS). Methodologically, it proposes a systematic evaluation framework integrating synthetic data generation, multi-judge verification, and psychometric analysis to conduct human-AI collaborative assessments of nine open-source LLMs. The findings reveal that model scale is not the sole determinant of performance, as certain smaller, locally deployable models demonstrate competitive results. Furthermore, human review proves essential for ensuring alignment with educational content. This research provides empirical evidence and methodological support for optimizing model selection and designing interactive feedback mechanisms in educational contexts.
📝 Abstract
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs'performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs'capacity for interactive, evidence-based feedback in educational scenarios.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Science Education
Benchmark
NGSS Alignment
K-12
Innovation

Methods, ideas, or system contributions that make the work stand out.

NGSS-aligned benchmark
synthetic data pipeline
multi-judge validation
psychometric analysis
human-in-the-loop
🔎 Similar Papers
No similar papers found.