Score
Designing and specifying question templates that elicit the intended capabilities (e.g., music perception rather than surface text cues) and generating diverse question families (continuous, binary, conditional) across arbitrary time horizons so evaluations require the targeted reasoning or forecasting behavior.
Low credibility, poor scalability, and ethical risks—including bias, privacy violations, and lack of transparency—hinder the use of generative AI, particularly large language models (LLMs), for autonomously generating high-quality, context-aware survey questions in engineering education research. Method: We propose a methodology framework for generative AI in educational surveys, introducing the first Synthetic Question–Response Analysis (SQRA) model grounded in Activity Theory to systematically model mediating mechanisms of learner engagement and associated ethical risks. Validation employs dual pathways—AI-to-AI and AI-to-human—via sentiment analysis, lexical statistics, and structured text analysis. Contribution/Results: Empirical findings demonstrate that prompt engineering combined with multidimensional validation significantly enhances question appropriateness, reliability, and validity. The study delineates the effective operational boundaries of AI-generated survey items and establishes a reusable, auditable, and scalable paradigm for intelligent data collection in educational empirical research.
Defining and generating pedagogically effective questions—those yielding measurable learning gains—remains a fundamental challenge in educational AI. Method: We propose QUEST, a framework that (1) formally defines and empirically estimates question utility based on real-world learning outcomes (e.g., post-intervention exam score improvements); (2) constructs an LLM-driven learning environment simulator to enable computationally tractable utility evaluation; and (3) introduces a utility-driven rejection sampling fine-tuning paradigm, replacing conventional approaches reliant on pedagogical heuristics or indirect proxies (e.g., information gain). Contribution/Results: Experiments demonstrate that QUEST-generated questions yield average exam score improvements exceeding 20%, significantly outperforming state-of-the-art baselines. This work establishes a novel, quantifiable, and optimization-friendly paradigm for learning-oriented question generation.
Existing automated math problem generation methods neglect pedagogical intent, support only unidimensional objectives, and suffer from misalignment between textual quality and educational appropriateness. To address this, we propose EQPR, an educational-goal-driven problem generation framework introducing the novel “Plan–Evaluate–Optimize” paradigm. Methodologically, EQPR integrates Monte Carlo Tree Search for education-goal-guided structured problem planning and leverages large language models for iterative self-reflective generation. Our contributions include: (1) EduMath—the first large-scale dataset comprising 16K problems annotated with fine-grained, three-dimensional educational goals (cognitive level, knowledge concept, and difficulty); and (2) EQGEVAL—a multidimensional alignment evaluation benchmark. Experiments demonstrate that EQPR significantly improves goal alignment on EQGEVAL and generates problems better aligned with instructional context and foundational pedagogical intentions.
Prior work lacks a systematic framework for assessing question quality. Method: This paper introduces the first dual-dimensional definition—question appropriateness (capturing sociolinguistic competence in context) and effectiveness (reflecting goal-directed strategic competence)—and builds a dynamically adaptive evaluation framework upon it. We design a semi-adaptive rubric-based scoring system integrating linguistic theory and dynamic contextual modeling, empirically validated on the CAUS and SQUARE datasets. Contribution/Results: The framework significantly improves discriminability between well-formed and defective questions, demonstrating robustness, interpretability, and contextual flexibility across diverse scenarios. Key contributions include: (1) establishing the first theoretically grounded definition of question quality; (2) proposing a scalable, principled assessment paradigm; and (3) providing a reusable foundational tool for evaluating AI question-asking capabilities.
This study addresses the low efficiency and uneven cognitive-level coverage in teacher-authored test item generation. We propose a novel test item auto-generation method integrating Bloom’s Taxonomy with large language models (LLMs), employing structured prompt engineering to explicitly encode Bloom’s six cognitive levels into the LLM generation process—enabling on-demand synthesis of multi-level, multi-format items. A rigorous empirical evaluation was conducted with frontline teachers, incorporating reliability analysis, item difficulty distribution assessment, and cognitive-level coverage evaluation. Results demonstrate that the generated items match or surpass human-authored items in reliability, difficulty appropriateness, and Bloom-level coverage; furthermore, teacher adoption intent is high, indicating strong potential for scalable classroom deployment. To our knowledge, this work represents the first effort to achieve deep, structured integration of Bloom’s Taxonomy into LLM-based test generation, accompanied by comprehensive pedagogical validation.
This study addresses the scarcity of high-quality, comprehensive, and secure multiple-choice questions in natural science education, which are costly to develop manually. The authors propose a human-AI collaborative two-stage generation framework: first, question prototypes are constructed using cognitive modeling templates; then, example-driven, single-step model expansion produces diverse item families targeting the same learning objective. The approach innovatively integrates multi-agent AI prompting with iterative human review to ensure generated items are not mere paraphrases and meet rigorous quality standards. Experimental results show that approximately 50% of the generated items are usable as-is, and minor revisions substantially increase the acceptance rate of entire item families while significantly improving prototype quality.
Current large language models struggle to reliably generate diverse multiple-choice questions aligned with specific cognitive levels—such as comprehension, reasoning, and main idea identification. To address this limitation, this work proposes the ReQUESTA framework, which introduces a novel hybrid multi-agent architecture that decomposes question generation into distinct phases: planning, constrained generation, iterative evaluation, and post-processing. By integrating large language models with rule-based engines and cognitive taxonomies, ReQUESTA enables structured and controllable generation of high-quality items. Experimental results demonstrate that the generated questions exhibit superior psychometric properties, including higher difficulty and discrimination indices, and show strong alignment with reading comprehension ability. Expert evaluations further confirm significant improvements over baseline methods in thematic relevance, distractor coherence, and semantic plausibility.
This study addresses the limitation of existing educational question generation methods in achieving fine-grained personalization and dynamically recommending exercises that maximize learning gains based on a student’s current knowledge state. To this end, the work proposes a novel personalized question recommendation framework that explicitly leverages knowledge tracing (KT) models to guide large language models (LLMs) in generating targeted practice questions. By predicting students’ mastery levels across knowledge concepts, the system identifies the most critical areas requiring reinforcement and produces corresponding exercises aimed at holistically improving overall knowledge acquisition. Experimental results on the XES3G5M and MOOCRadar datasets demonstrate that the generated questions significantly outperform baseline approaches lacking or offering only limited personalization in terms of instructional effectiveness.
This work addresses the limitation of existing large language models in psychotherapy, which typically operate passively without actively guiding cognitive restructuring. The authors propose the Socratic Inquiry Framework (SIF), which for the first time decouples theory-driven Socratic questioning into a plug-and-play module. SIF employs strategy anchoring to determine optimal questioning时机, retrieves templates to generate question content, and integrates a lightweight intent-planning architecture—enabling proactive therapeutic guidance without requiring end-to-end retraining. A high-quality, strategy-aligned Socratic-QA dataset is also introduced. Experimental results demonstrate that SIF significantly increases the frequency of proactive questioning, enhances dialogue depth, and improves therapeutic alignment, thereby achieving a paradigm shift from passive empathy to active cognitive guidance.
This study addresses the tendency of users to uncritically accept recommendations from AI decision-support systems, which can lead to erroneous judgments. To mitigate this issue, the authors propose a data-driven prompting mechanism that enhances prospective reasoning in human-AI collaboration by automatically generating reflective questions via large language models (LLMs). The approach integrates a structured question taxonomy, a cognitive engagement scale, and LLM-based question generation. A prototype system was developed and evaluated in a clinical setting, combining methods from LLMs, human-computer interaction design, and cognitive assessment. Empirical results demonstrate that the proposed mechanism significantly improves clinicians’ critical appraisal of AI-generated outputs, eliciting positive user feedback and offering a novel pathway toward developing AI systems that function as “thinking tools” rather than mere decision aids.