Score
Designs and conducts semi-structured qualitative interviews and the associated recruitment and informed-consent procedures; facilitates in-depth participant conversations, transcribes and codes interview data, synthesizes thematic findings, and interprets participant narratives into actionable insights.
In semi-structured interviews, the quality of follow-up questioning is highly dependent on interviewer expertise, and the potential of large language models (LLMs) to augment data collection remains underexplored. Method: We introduce the “AI-augmented puppeteer” paradigm, embedding an LLM into real-time interview workflows via a Wizard-of-Oz experimental design to generate context-sensitive follow-up questions, and systematically examine human–AI dynamics in role allocation, collaborative behavior, and responsibility distribution. Based on an empirical study with 17 participants, we develop a human–AI co-interviewing framework and human-centered design guidelines. Results: Findings confirm that LLMs significantly enhance the depth and topical breadth of follow-up questions—but only when humans retain epistemic authority and ethical oversight. Our core contribution is the first empirical demonstration of how LLMs function as *collaborators*—not substitutes—in qualitative data collection, revealing their impact mechanism on data quality and establishing a methodological foundation and practical pathway for AI-enhanced qualitative research.
This study addresses the challenge interviewers face in semi-structured interviews, where high cognitive load impedes their ability to simultaneously listen attentively, adapt interview guides flexibly, and pose effective follow-up questions. To mitigate this, the authors propose InterFlow—a non-intrusive, user-directed AI-assisted system that automates three-tiered information capture through dynamic script adaptation, visual progress tracking, AI-generated summaries, and a collaborative agent. The system also allows users to specify focal points of interest. In a within-subjects experiment with twelve participants, InterFlow significantly reduced cognitive load while enhancing both interview efficiency and user experience, offering a novel paradigm for human-AI collaborative decision-making in high-stakes, time-sensitive scenarios.
Existing LLM-based interview systems struggle to balance predefined topic coverage with adaptive exploration, limiting the scalable acquisition of high-quality qualitative user insights. This work proposes a multi-agent LLM architecture that frames adaptive semi-structured interviewing as a utility optimization problem, formally defining interview utility as a trade-off among topic coverage, discovery of novel insights, and conversational cost. The system dynamically plans high-expected-utility questions through simulated dialogue rollouts. Experiments demonstrate that, in LLM simulations, the approach improves topic coverage by 4.7% and yields richer insights in fewer turns. A user study with 70 participants further validates that domain experts recognize the method’s ability to uncover high-quality insights in professional contexts that existing approaches fail to capture.
In participatory AI, stakeholder recruitment continues to face challenges—including identification bias, access barriers, and insufficient inclusivity—that undermine equity and empowerment goals. This study systematically examines structural limitations in recruitment practices through a literature review of 37 AI projects and in-depth interviews with 5 researchers, analyzed via qualitative content analysis. We introduce a novel “relationship-first” recruitment framework that foregrounds the dynamic interplay among structural conditions, researcher intent, and collaborative relationships, and propose reflective recruitment documentation standards. The work clarifies how recruitment practices fundamentally shape participation quality and offers actionable, relationship-centered design principles and implementation guidelines. By centering relationality and reflexivity, this research advances a methodological foundation for more inclusive and empowering participatory AI.
Traditional survey methods face a trade-off between depth and scalability: structured questionnaires scale well but lack expressive flexibility, whereas in-depth interviews yield rich insights yet are labor-intensive and difficult to scale. Method: This study conducts the first controlled experimental evaluation of large language models (LLMs) as adaptive, conversational interviewers—specifically for political topics—comparing AI- and human-administered interviews across data quality, participant engagement, and operational efficiency. We propose a design framework that reconciles standardization with conversational adaptability, integrating structured questionnaire logic, real-time response generation, and multi-dimensional evaluation metrics (e.g., protocol adherence, response quality, engagement). Contribution/Results: AI-conducted interviews achieve data quality comparable to human interviews, demonstrate substantially improved scalability, and elicit positive participant feedback. This work establishes a novel paradigm for high-fidelity, large-scale qualitative data collection in the social sciences.
This study addresses the challenges of traditional semi-structured interviews in empirical software engineering, which are often resource-intensive and hindered by cross-time-zone coordination and multilingual barriers. The authors propose a self-administered AI interview approach based on a customized MyGPT model, enabling participants to complete unmoderated interviews via voice in their preferred language, with the system automatically generating structured summaries according to a predefined protocol. As the first work to demonstrate the feasibility of AI-conducted, short-duration, low-risk interviews in this domain, the evaluation shows that 92.4% of 66 submissions met formatting requirements; 90.9% of participants reported a positive experience, 95.5% found the questions clear, and 89.4% expressed willingness to participate again, indicating high acceptability and effectiveness. The study also identifies limitations concerning interview depth and privacy concerns.
This study addresses the challenges interviewers face in requirements elicitation interviews, where balancing comprehensive topic coverage, active listening, and adaptive follow-up questioning is difficult, compounded by a lack of effective script execution tracking. To overcome these limitations, this work proposes the first end-to-end AI-assisted interview framework, integrating business-goal-driven theoretical script generation, real-time topic coverage monitoring via natural language processing, and an on-demand dynamic probing mechanism. Experimental results demonstrate that the proposed approach significantly improves script quality (92.8 vs. 74.8), probing depth (3.43 vs. 1.15 probes per topic), and granularity of the resulting requirements models (proportion of leaf-level goals: 0.653 vs. 0.598). Notably, 86% of users identified real-time topic tracking as the most practically valuable feature.
Traditional questionnaires struggle to simultaneously capture qualitative depth and quantitative structure, limiting comprehensive understanding of complex social phenomena. This study proposes a dynamic survey platform powered by large language models (LLMs) that, for the first time, enables real-time semantic clustering of open-ended responses. Through an interactive feedback mechanism, users can rate, rank, and reflect on these clusters, generating visual reports that integrate qualitative insights with quantitative analysis. Innovatively embedding LLMs within a closed-loop data collection framework, the approach facilitates dynamic comparisons between individual perspectives and group-level trends. Empirical validation across two field studies involving 93 participants demonstrates that the platform significantly enhances data richness and user engagement compared to conventional survey tools, while effectively fostering collaborative sensemaking.
This study addresses the high labor costs of semi-structured interviews, the opaque behavioral mechanisms of existing multimodal large language models (MLLMs), and the challenges in establishing trust during human–AI interaction. To this end, the authors developed InterviewBot—a system that integrates researcher-designed interview protocols—and conducted the first empirical analysis of MLLM turn-by-turn behaviors in a real-world deployment setting. Leveraging voice-driven real-time MLLM interaction, a semi-structured interview framework, and inductive qualitative methods, the study identifies four data collection failure modes, including information loss and premature termination, and reveals that only 4.9% of AI-generated questions constituted probing follow-ups while 28.7% violated single-question instructions. The findings uncover three key socio-dynamic mechanisms—disclosure calibration, institutional legitimacy, and conversational anchoring—and distill design principles centered on enhanced depth control and non-scripted listening.