CLARA: Can AI Assess Developmental Appropriateness in Children's Stories?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on manual evaluation, scalability limitations, and inadequate AI alignment in assessing the developmental appropriateness of children's stories by proposing a cognition-oriented multidimensional evaluation framework. Methodologically, we construct a three-dimensional structured annotation schema encompassing cognitive, linguistic, and socio-emotional dimensions, and establish a bilingual benchmark dataset that transcends the constraints of traditional readability metrics. Furthermore, we conduct systematic evaluations by integrating large language model prompt engineering with multidimensional component analysis. Experimental results demonstrate that the proposed approach significantly outperforms baselines and effectively approximates human expert judgment, thereby validating the potential of explainable AI for educational natural language processing applications.
📝 Abstract
Assessing the developmental suitability of children's narratives is important for educational recommendation and developmental literacy research, yet such assessment typically relies on subjective and difficult-to-scale human judgment. This raises an important question: Can AI systems approximate human developmental judgments of children's stories? To study this problem, we introduce CLARA, a cognitively grounded framework for developmental narrative understanding through structured annotation across cognitive (COG), language (LAN), and social-emotional (SEL) dimensions, together with a bilingual benchmark resource containing 1107 Chinese--English children's stories with normalized silver developmental references and structured developmental annotations. We evaluate CLARA through benchmark comparison, component analysis, translated bilingual consistency analysis, and blinded human evaluation with educators. Experimental results show that structured developmental annotation achieves substantially stronger alignment with developmental references and human judgments than readability-based methods and direct prompting baselines. Overall, our findings suggest that AI systems can approximate certain aspects of human developmental judgment when guided by structured developmental annotation, while also highlighting the importance of interpretability and human oversight in educational NLP.
Problem

Research questions and friction points this paper is trying to address.

developmental appropriateness
children's stories
AI assessment
educational NLP
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developmental Appropriateness
Structured Annotation
Cognitively Grounded Framework
Bilingual Benchmark
Educational NLP
💼 Related Jobs
No related jobs found.