Score
Designing and conducting semi-structured interviews and complementary survey methods to elicit recurring forms, causes, and stakeholders’ sense-making from field settings. Used to gather qualitative evidence about adaptations, interpretations of monitoring outputs, and coordination behaviors across issue lifecycles.
Ambiguous definitions of Participatory Design (PD) have led to conceptual vagueness and unresolved concerns regarding design fairness. Method: We conducted a systematic literature review (SLR) of over 100 empirical PD studies, applying thematic coding and cross-case comparison. Contribution/Results: First, we identify—structurally and for the first time—five core leverage points (e.g., emergent vs. pre-specified design, direct vs. indirect participation) that mediate the relationship between PD processes and fairness outcomes, thereby establishing a theoretical framework linking PD practice to design fairness. Second, we catalog 14 concrete participatory techniques, revealing intangible system design as the dominant application domain and multi-stage recruitment with hybrid technique combinations as prevailing practices. Third, we clarify how stakeholders’ degree, timing, mode, and technical configuration of involvement shape fairness mechanisms. This work provides empirically grounded, actionable decision guidelines for advancing PD methodology and practice.
In semi-structured interviews, the quality of follow-up questioning is highly dependent on interviewer expertise, and the potential of large language models (LLMs) to augment data collection remains underexplored. Method: We introduce the “AI-augmented puppeteer” paradigm, embedding an LLM into real-time interview workflows via a Wizard-of-Oz experimental design to generate context-sensitive follow-up questions, and systematically examine human–AI dynamics in role allocation, collaborative behavior, and responsibility distribution. Based on an empirical study with 17 participants, we develop a human–AI co-interviewing framework and human-centered design guidelines. Results: Findings confirm that LLMs significantly enhance the depth and topical breadth of follow-up questions—but only when humans retain epistemic authority and ethical oversight. Our core contribution is the first empirical demonstration of how LLMs function as *collaborators*—not substitutes—in qualitative data collection, revealing their impact mechanism on data quality and establishing a methodological foundation and practical pathway for AI-enhanced qualitative research.
Existing LLM-based interview systems struggle to balance predefined topic coverage with adaptive exploration, limiting the scalable acquisition of high-quality qualitative user insights. This work proposes a multi-agent LLM architecture that frames adaptive semi-structured interviewing as a utility optimization problem, formally defining interview utility as a trade-off among topic coverage, discovery of novel insights, and conversational cost. The system dynamically plans high-expected-utility questions through simulated dialogue rollouts. Experiments demonstrate that, in LLM simulations, the approach improves topic coverage by 4.7% and yields richer insights in fewer turns. A user study with 70 participants further validates that domain experts recognize the method’s ability to uncover high-quality insights in professional contexts that existing approaches fail to capture.
This study addresses the structural misalignment between qualitative and quantitative data in mixed-methods research by leveraging large language models (LLMs) to generate psychometrically reliable synthetic survey responses from interview transcripts. Methodologically, we employed the Behavioral Regulation in Exercise Questionnaire (BREQ) as the measurement framework, integrating content from in-depth interviews with after-school program staff. Using structured prompt engineering and low-temperature sampling, we systematically evaluated—across Claude and GPT models—how interview-derived contextual cues enhance response quality. Key contributions include: (1) the first empirical demonstration that interview guidance significantly improves both response diversity and fidelity to individual response patterns; (2) evidence that prompt design and temperature parameters exert stronger influence on psychometric alignment than demographic variables; and (3) confirmation that while LLMs reliably reproduce aggregate distributions, they initially underrepresent response variability—a limitation effectively mitigated by interview-informed conditioning.
In participatory AI, stakeholder recruitment continues to face challenges—including identification bias, access barriers, and insufficient inclusivity—that undermine equity and empowerment goals. This study systematically examines structural limitations in recruitment practices through a literature review of 37 AI projects and in-depth interviews with 5 researchers, analyzed via qualitative content analysis. We introduce a novel “relationship-first” recruitment framework that foregrounds the dynamic interplay among structural conditions, researcher intent, and collaborative relationships, and propose reflective recruitment documentation standards. The work clarifies how recruitment practices fundamentally shape participation quality and offers actionable, relationship-centered design principles and implementation guidelines. By centering relationality and reflexivity, this research advances a methodological foundation for more inclusive and empowering participatory AI.
Traditional survey methods face a trade-off between depth and scalability: structured questionnaires scale well but lack expressive flexibility, whereas in-depth interviews yield rich insights yet are labor-intensive and difficult to scale. Method: This study conducts the first controlled experimental evaluation of large language models (LLMs) as adaptive, conversational interviewers—specifically for political topics—comparing AI- and human-administered interviews across data quality, participant engagement, and operational efficiency. We propose a design framework that reconciles standardization with conversational adaptability, integrating structured questionnaire logic, real-time response generation, and multi-dimensional evaluation metrics (e.g., protocol adherence, response quality, engagement). Contribution/Results: AI-conducted interviews achieve data quality comparable to human interviews, demonstrate substantially improved scalability, and elicit positive participant feedback. This work establishes a novel paradigm for high-fidelity, large-scale qualitative data collection in the social sciences.
This study addresses the absence of empirically validated quality metrics for evaluating the contribution of interview responses to qualitative research objectives. Building a corpus of 343 interview transcripts comprising 16,940 responses, the authors systematically assess the predictive validity of ten established quality indicators with respect to their research utility. Integrating qualitative content analysis, natural language processing, and statistical modeling, the findings demonstrate that direct relevance to the core research question is the strongest predictor of a response’s value, whereas commonly used NLP-based metrics—such as clarity and surprise-based informativeness—show no significant predictive power. These results challenge the applicability of current automated evaluation approaches and provide empirical grounding for assessing response quality in qualitative inquiry.
This study addresses the lack of standardized reporting criteria and operational definitions for sample size adequacy and data saturation in software engineering interview research. Through a systematic analysis of 138 interview-based studies published in major software engineering conferences between 2016 and 2025, the authors employ bibliometric and qualitative content analysis methods to code and synthesize reported sample sizes, discussions of saturation, and justifications for methodological choices. This work presents the first large-scale empirical mapping of operational practices regarding interview sample sizes in the field, revealing that most studies recruit between 13 and 24 participants, while samples smaller than 12 are typically confined to specific industrial contexts. Notably, the majority of studies fail to explicitly articulate their criteria for determining data saturation. These findings contribute to enhancing transparency and methodological rigor in qualitative software engineering research.
This study addresses the challenges of traditional semi-structured interviews in empirical software engineering, which are often resource-intensive and hindered by cross-time-zone coordination and multilingual barriers. The authors propose a self-administered AI interview approach based on a customized MyGPT model, enabling participants to complete unmoderated interviews via voice in their preferred language, with the system automatically generating structured summaries according to a predefined protocol. As the first work to demonstrate the feasibility of AI-conducted, short-duration, low-risk interviews in this domain, the evaluation shows that 92.4% of 66 submissions met formatting requirements; 90.9% of participants reported a positive experience, 95.5% found the questions clear, and 89.4% expressed willingness to participate again, indicating high acceptability and effectiveness. The study also identifies limitations concerning interview depth and privacy concerns.
This study addresses measurement error introduced when mapping natural language responses to structured variables in AI-assisted interviews, particularly noting its heterogeneous impact across subpopulations. The authors propose the Adaptive Matrix Validation (AMV) framework, which employs a small random subset of structured validation questions and integrates both cross-respondent calibration and within-respondent verification to doubly correct AI-derived mappings. AMV is the first approach to jointly leverage sparse validation data and inter-respondent calibration, enabling unbiased inference for population means, subgroup parameters, and regression coefficients. The framework also provides a joint planning formula for determining the required number of validation items and sample size. Empirical evaluations—including design-based simulations, a replication using the American Time Use Survey, and an application to CHAMPS verbal autopsy data—demonstrate that AMV substantially improves estimation accuracy even with limited validation resources.
This study addresses the lack of systematic understanding regarding the types and motivations of visual representations in qualitative research. Building upon and extending Verdinelli & Scagnoli’s (2013) work through a data-driven literature review, it conducts a content analysis of articles and their visualizations published between 2020 and 2022 in three leading qualitative methods journals. Integrating epistemological stance classification with visualization-type coding, the study innovatively combines correspondence analysis and cognitive network analysis for the first time. Findings indicate that while visualizations remain underutilized in qualitative research, their typological diversity is increasing, and the choice of graphical representation appears largely independent of the authors’ epistemological positions. These results offer both empirical grounding and methodological innovation for integrating interdisciplinary visualization tools into qualitative inquiry.