🤖 AI Summary
This study addresses the scalability challenges of traditional manual coding when analyzing large-scale textual data generated from K–12 teachers’ interactions with AI systems. It proposes a human-AI collaborative qualitative analysis framework that positions large language models (LLMs) as assistive annotation tools rather than interpretive agents, while preserving human researchers’ central role in defining concepts and constructing analytical frameworks through open, axial, and selective coding. Key innovations include an auditable human-AI workflow, a set-valued intercoder agreement metric suited for multi-label contexts, and a multi-round calibration mechanism. The resulting codebook comprises 72 codes organized into 19 categories across six domains, demonstrating high reliability on a dataset of 2,560 messages. The approach’s validity and scalability are further evidenced by human researchers identifying five novel codes missed by the LLM.
📝 Abstract
Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educational domain expertise applied the codebook to an independent sample of 2,560 messages, established reliability through iterative calibration using set-valued agreement measures appropriate for multi-label annotation, and extended the instrument with five codes that the LLM-assisted phases had not surfaced. The final codebook comprises 72 items within 19 categories and six domains. We reflect on the methodological decisions the pipeline required, including the choice of a conversational unit of analysis, the treatment of the LLM as a labeling instrument rather than an interpretive agent, the measurement of intercoder agreement under multi-label coding, and the conditions under which human domain expertise remained decisive. The account is offered as an auditable template for qualitative researchers considering LLM assistance in codebook development while preserving human interpretive authority.