Score
Applying qualitative research methods (coding, thematic synthesis, expert comparison) to interpret human judgments, deployment effects, and recurring patterns in nonnumeric data and to surface failure modes and adaptation behaviours.
Software engineering (SE) practices are often too complex for quantitative methods to fully capture, limiting empirical understanding of developer behavior, team collaboration, and organizational contexts. To address this, this study conducts multiple expert focus groups and applies qualitative content analysis and thematic synthesis—yielding the first structured, dialogic account of qualitative SE research’s current state and future trajectory. It innovatively positions “narrative” as a core epistemological resource in empirical SE, underscoring the irreplaceable role of qualitative inquiry in uncovering situated, processual, and socially embedded phenomena. The study identifies three key pathways for advancement: methodological integration (e.g., mixed-methods designs), systematic training infrastructure for qualitative literacy, and sustained cross-paradigmatic dialogue between positivist and interpretivist traditions. These contributions provide both theoretical grounding and actionable guidance for fostering pluralistic, rigorous, and context-sensitive empirical research in SE. (149 words)
This study addresses key limitations in gerontological qualitative research—namely, constraints in data scale, pattern detection, and methodological integration. Methodologically, it pioneers the embedding of machine learning and natural language processing (NLP) techniques directly into qualitative workflows, enabling systematic indexing, multi-scale textual analysis, and reproducible management of participatory observation and in-depth interview data. It integrates open science platforms with qualitative data management systems to support synergistic analysis of large-scale secondary datasets (e.g., the American Voices Project) and original ethnographic fieldwork (e.g., DISCERN dementia ethnography). Contributions include: (1) substantially enhancing qualitative data processing efficiency and analytical transparency; (2) achieving a principled integration of humanistic depth with computational breadth; and (3) advancing aging research toward a multimodal, scalable, and verifiable mixed-methods paradigm.
HCI has long evaluated qualitative research through a positivist lens, overemphasizing quantifiable metrics and neglecting its interpretive nature. Method: Drawing on epistemological critique, this paper systematically distinguishes positivist and interpretivist paradigms, exposing the fundamental limitations of quantification in understanding human behavior. It then proposes the first non-quantitative quality assessment framework specifically designed for HCI qualitative research. Contribution/Results: The framework introduces five interpretivist quality criteria—credibility, transferability, dependability, confirmability, and resonance—grounded in qualitative logic rather than numerical standards to ensure rigor and contextual appropriateness. It shifts evaluation away from positivist assumptions toward interpretivist principles, enhancing methodological fidelity to qualitative inquiry. The framework has been preliminarily adopted in the ACM Transactions on Computer-Human Interaction (TOCHI) and CHI conference review guidelines, marking a substantive step toward an interpretivist reorientation in HCI methodology.
This study systematically investigates the applicability boundaries and risks of misuse of generative artificial intelligence (GenAI) in qualitative software engineering research, emphasizing that GenAI is not a universal solution. By examining the multidimensional nature of qualitative inquiry and accounting for diverse research strategies and data characteristics, the work integrates large language model applications, qualitative data analysis, and empirical evidence to reconstruct quality assessment criteria. It elucidates the critical mechanisms governing the alignment between technological capabilities and methodological requirements. The research delineates specific scenarios where GenAI offers advantages as well as contexts prone to pitfalls, proposes principled guidelines for appropriate adoption, and offers practical guidance for researchers while charting directions for future investigation.
Qualitative evaluation of inductive coding faces significant challenges: conventional metrics are ill-suited for exploratory processes; manual assessment is labor-intensive; and expert-annotated “ground truth” introduces methodological limitations. This paper introduces the first quantifiable evaluation framework for open coding in grounded theory and thematic analysis. Moving beyond the conventional human–AI alignment paradigm, it innovatively integrates stability assessment (Cohen’s kappa, Jaccard similarity) with cross-human–machine comparison as a dual-validation mechanism. Crucially, the framework operates without presupposing ground-truth labels, enabling bias detection and coding quality measurement in human–AI collaborative workflows. Evaluated on two HCI datasets, the framework demonstrates high inter-coder agreement (κ > 0.75), strong output stability (Jaccard similarity > 0.89 across repeated runs), and yields a reusable, AI-augmented coding workflow.
Existing computational tools for qualitative data analysis often fall short in effectively supporting causal exploration due to insufficient contextual awareness, limited trustworthiness, or overly complex outputs. To address these limitations, this work proposes QualCausal, the first interactive causal analysis system grounded in user research–driven design principles. Developed through formative user studies, QualCausal integrates context-aware processing, cognitive scaffolding, and explainability mechanisms to facilitate efficient exploration and validation of causal hypotheses within qualitative datasets. The system enables researchers to extract causal relationships, construct interactive causal networks, and examine findings through coordinated multi-view visualizations. User evaluations demonstrate that QualCausal significantly reduces analytical burden, provides robust cognitive support, and prompts critical reflection on how computational tools can be meaningfully integrated into social science research practices, thereby bridging the gap between computational assistance and qualitative inquiry paradigms.
This study examines the applicability and contested boundaries of generative AI in qualitative research. It innovatively distinguishes between “small-q” positivist and “big-Q” non-positivist paradigms, using this dichotomy as a central criterion to develop a multidimensional decision framework for AI adoption that incorporates researcher expertise, ethical considerations, and personal preferences. Through a comprehensive literature review and theoretical analysis, the paper clarifies the conditions under which generative AI is methodologically justifiable or constrained within each paradigm. The findings offer valuable methodological guidance and practical insights for conducting qualitative research—particularly in fields such as software engineering—where the integration of AI tools raises both opportunities and epistemological challenges.
This study investigates the reliability of large language models (LLMs) in replicating human reasoning for qualitative coding of psychological safety in software engineering communities, with a focus on performance disparities and systematic biases across different prompting strategies. Through controlled experiments, the authors evaluate Cohen’s κ agreement and stability of Claude Haiku, DeepSeek-Chat, and Gemini 2.5 Flash under zero-shot and few-shot closed-ended prompting. The work presents the first systematic quantification of few-shot prompting effects on LLM-based qualitative coding, revealing that this strategy significantly improves Claude Haiku’s intercoder agreement (Δκ = +0.034). Claude Haiku and DeepSeek-Chat demonstrate the highest stability (SD ≈ 0.017). All models consistently over-predict “sharing negative feedback” and under-predict “expressing concerns,” offering empirical insights and methodological guidance for LLM-assisted qualitative research.
This study addresses the lack of systematic understanding regarding the types and motivations of visual representations in qualitative research. Building upon and extending Verdinelli & Scagnoli’s (2013) work through a data-driven literature review, it conducts a content analysis of articles and their visualizations published between 2020 and 2022 in three leading qualitative methods journals. Integrating epistemological stance classification with visualization-type coding, the study innovatively combines correspondence analysis and cognitive network analysis for the first time. Findings indicate that while visualizations remain underutilized in qualitative research, their typological diversity is increasing, and the choice of graphical representation appears largely independent of the authors’ epistemological positions. These results offer both empirical grounding and methodological innovation for integrating interdisciplinary visualization tools into qualitative inquiry.