CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health Research

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data fragmentation and privacy constraints in childcare quality research by introducing a novel, large language model (LLM)-based automated de-identification pipeline. Integrating nearly 60,000 records across twelve states, we construct a large-scale dataset that balances privacy preservation with research utility, provided in both cleaned textual and preprocessed tabular formats. Methodologically, this work combines LLM-assisted data cleaning, traditional machine learning, tabular foundation models, and cross-domain transfer learning. Experimental results demonstrate that supervised fine-tuning significantly outperforms zero-shot transfer. To accelerate AI-driven early childhood health research, all datasets and source code have been made publicly available.
📝 Abstract
High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a large-scale, de-identified dataset for applied data science research at the intersection of AI and early childhood health. CCQ integrates 59,372 child care provider records across 12 U.S. states, covering diverse provider types as well as data schemas. To ensure research utility while protecting privacy, we implement an automated, LLM-based curation pipeline that anonymizes, cleans, and standardizes raw state records into two complementary releases: a cleaned textual release and a fully preprocessed tabular release. We also benchmark traditional machine learning models, tabular foundation models, and language models on quality rating prediction and important features analytics. Within a state, tabular classifiers on the preprocessed tables perform best. Across states, zero-shot transfer is near chance, but modest target-state supervision recovers most of the within-state performance, and pretraining on other states benefits finetuned language models. We release both datasets with all code to accelerate AI-driven research on child care quality and ultimately improve children's health and development.
Problem

Research questions and friction points this paper is trying to address.

child care quality
dataset fragmentation
privacy constraints
children's health
AI research
Innovation

Methods, ideas, or system contributions that make the work stand out.

Child Care Quality Dataset
LLM-based Curation Pipeline
Tabular Foundation Models
Cross-State Transfer Learning
Data Anonymization
💼 Related Jobs
No related jobs found.
Victor Li
Victor Li
Emory University, USA
Y
Yuzhang Xie
Emory University, USA
Z
Ziwei Dong
Emory University, USA
Q
Qingyang Zhu
Emory University, USA
W
Wenjing Ma
Emory University, USA
Carl Yang
Carl Yang
Waymo LLC, PhD at University of California, Davis
GPU ComputingParallel ComputingGraph Processing
J
Jinbing Bai
Emory University, USA
H
Huiwen Xu
Emory University, USA
Jiaying Lu
Jiaying Lu
Research Assistant Professor of School of Nursing's Center for Data Science, at Emory University
AI for HealthcareKnowledge GraphMultimodal LearningLarge Language Model