Score
Designs and analyzes artifacts that map learning units, courses, or curricula to stated competencies and quantify the cognitive depth at which each competency is taught. Builds metrics, measurement processes, and reports (e.g., coverage maps, depth scores, articulation-gap analyses) that compare delivered cognitive depth to recommended targets and identify gaps across versions.
This study addresses the deployment challenges of cognitive diagnosis models caused by their reliance on manually constructed Q-matrices by proposing an automated framework for course knowledge graph generation. The method employs a dual-repository system and a multi-agent collaborative pipeline to automatically extract and validate concept-skill mappings from official documentation and textbooks, shifting the judgment cost from the item level to the course level. Furthermore, it introduces an eleven single-responsibility agent architecture featuring self-healing and human-escalation mechanisms, a deterministic orchestrator, and multi-source cross-validation techniques. Experimental results demonstrate that across 241 runs, the system achieves 57.0% internal consistency after 77 rounds, increases the non-intervention rate to 91.5%, and incurs a processing cost of merely $1.19 per chapter.
This study addresses the lack of reliable methods for evaluating how undergraduate computer science curricula align with international teaching guidelines such as CS2013 and CS2023, and how this alignment evolves across guideline revisions. The authors propose a human-in-the-loop analytical pipeline that structures course content and guideline knowledge units, employs semantic retrieval—incorporating reciprocal rank fusion and lightweight sentence embedding models—to generate matching candidates, and applies clearly defined coverage criteria for human validation, supplemented by Cohen’s kappa to assess inter-rater consistency. For the first time, the approach enables longitudinal measurement of curriculum alignment across three dimensions: topic coverage, competency expression, and cognitive depth, distinguishing structural gaps from changes due to updated standards. Empirical results reveal knowledge unit coverage rates of 50.9% for CS2013 and 49.7% for CS2023, competency coverage around 88%, but a notable decline in adherence to recommended cognitive depth—from 95% to 76%.
This study investigates the extent to which higher education curricula embed 21st-century core competencies and align with societal demands. To this end, we construct a dataset comprising 7,600 human-annotated samples and introduce a "Curricular Chain-of-Thought" (Curricular CoT) prompting strategy to enhance large language models’ reasoning capabilities in educational contexts, mitigate keyword-matching biases, and improve detection of subtle pedagogical evidence in lengthy texts. Experimental results demonstrate that detailed descriptions of teaching activities are the most informative; open-source models perform comparably to closed-source counterparts on coarse-grained competency mapping tasks; and while Curricular CoT yields a modest yet statistically significant performance gain, models still fall substantially short of human-level proficiency in fine-grained educational reasoning.
Intelligent tutoring systems (ITS) in curriculum-based online learning risk exacerbating academic achievement gaps among students. Method: This paper proposes CTGraph, the first self-supervised graph representation learning framework that explicitly incorporates curriculum structure priors to construct student behavior graphs—modeling multidimensional learning signals including learning pathways, content coverage, engagement intensity, and conceptual mastery. A graph neural network performs graph-level encoding to enable cross-cohort behavioral comparison and stage-wise difficulty localization. Contribution/Results: Experiments demonstrate that CTGraph accurately identifies at-risk students, pinpoints optimal intervention timing with fine-grained temporal resolution, and localizes specific knowledge gaps. It significantly enhances personalized instructional support while providing interpretable, pedagogically grounded insights—offering a transparent, equity-oriented technical pathway for adaptive education.
This study addresses the limitations of existing visualization literacy assessments, which predominantly rely on multiple-choice questions and struggle to measure higher-order competencies due to ceiling effects. To overcome this, the work introduces two web-based qualitative assessment methods—visualization critique and sketching tasks—implemented through online think-aloud protocols and data-driven drawing exercises, respectively. These approaches holistically capture users’ advanced literacy in interpreting, evaluating, and constructing visualizations. Leveraging interactive tools, controlled comparative experiments, and validity evidence from established scales such as CALVI and Mini-VLAT, the proposed methods significantly differentiate individuals across varying expertise levels—including researchers, students, and crowdworkers—thereby surpassing the constraints of traditional item formats and uncovering critical competency dimensions previously unmeasurable by existing instruments.
研究通过分析课程目标和教师反馈,提出可视化设计的概念清单,以明确学生应掌握的关键概念和技能。
This study addresses the challenge of interpreting the relationship between programming assignment difficulty and student performance, a gap exacerbated by the limited interpretability of existing predictive models that hinders effective pedagogical refinement. To bridge this gap, the authors propose an interpretable analytical framework grounded in knowledge components (KCs), which can be either expert-defined or automatically extracted using large language models (LLMs). By quantifying the number of KCs involved in each assignment and measuring the shift in KCs between consecutive tasks, the framework elucidates how assignment design influences learning outcomes. Empirical evaluation across three introductory programming course datasets reveals that assignments involving a greater number of KCs correlate with poorer student performance, and abrupt KC transitions are significantly associated with learning disruptions. These findings enable the identification of poorly designed assignments, offering actionable insights for instructional diagnosis and improvement.
研究探讨了模型素养作为视觉分析性能的额外评价因素,通过控制实验发现模型-任务准确性和视觉分析任务准确性之间存在正相关关系。
研究通过大型语言模型自动生成和评估包含错误的案例问题,自动评分并提供反馈,以解决HDR方法在评估实践技能时需要专业知识和大量努力的问题。
This study addresses a critical limitation in current learning analytics tools, wherein frequency-oriented visualizations often obscure rare yet educationally significant student feedback. To bridge the gap between quantitative visualization and qualitative educational research, the authors engaged STEM education researchers in analyzing student logs using the WordStream platform. Through an integrated approach combining thematic analysis, member checking, and mixed-methods user research, the study uncovered epistemological tensions educators face when repurposing quantitative codings for qualitative inquiry. Three core themes emerged: tool experience, disciplinary contextualization, and the integration of quantitative and qualitative paradigms. Building on these insights, the work proposes design principles for visualizations that support deep qualitative exploration, offering both theoretical grounding and practical guidance for the next generation of learning analytics tools.