🤖 AI Summary
Qualitative research faces significant bottlenecks in manual transcription—low efficiency, high time cost, and challenges in GDPR compliance and software interoperability. This study proposes an end-to-end AI-assisted transcription workflow: it employs a localized automatic speech recognition (ASR) model, augmented with customized text post-processing and structured format conversion modules. Crucially, it achieves native, seamless integration with leading qualitative data analysis software—including ATLAS.ti and MAXQDA—for direct import of transcripts. The system incorporates built-in GDPR compliance via fully offline processing and includes phonetic and linguistic adaptations for non-native speaker speech. Empirical evaluation across 12 real-world interviews demonstrates a 46.2% average reduction in transcription time, while preserving transcription accuracy, data privacy, and cross-platform compatibility. This workflow significantly enhances the efficiency and rigor of qualitative data preparation.
📝 Abstract
In qualitative research, data transcription is often labor-intensive and time-consuming. To expedite this process, a workflow utilizing artificial intelligence (AI) was developed. This workflow not only enhances transcription speed but also addresses the issue of AI-generated transcripts often lacking compatibility with standard content analysis software. Within this workflow, automatic speech recognition is employed to create initial transcripts from audio recordings, which are then formatted to be compatible with content analysis software such as ATLAS.ti or MAXQDA. Empirical data from a study of 12 interviews suggests that this workflow can reduce transcription time by up to 46.2%. Furthermore, by using widely used standard software, this process is suitable for both students and researchers while also being adaptable to a variety of learning, teaching, and research environments. It is also particularly beneficial for non-native speakers. In addition, the workflow is GDPR-compliant and facilitates local, offline transcript generation, which is crucial when dealing with sensitive data.