ai-assisted interviewing

Designs, builds, and evaluates interview workflows, interfaces, and evaluation methods that integrate LLMs into live qualitative interviews to generate context-sensitive follow-up questions, surface and enable real‑time editing or relaying of AI suggestions, and preserve human interviewer oversight. Also develops and runs simulation and wizard-of-oz setups to test interaction patterns, safety measures, and deployment protocols before live use.

ai-assistedinterviewing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In semi-structured interviews, the quality of follow-up questioning is highly dependent on interviewer expertise, and the potential of large language models (LLMs) to augment data collection remains underexplored. Method: We introduce the “AI-augmented puppeteer” paradigm, embedding an LLM into real-time interview workflows via a Wizard-of-Oz experimental design to generate context-sensitive follow-up questions, and systematically examine human–AI dynamics in role allocation, collaborative behavior, and responsibility distribution. Based on an empirical study with 17 participants, we develop a human–AI co-interviewing framework and human-centered design guidelines. Results: Findings confirm that LLMs significantly enhance the depth and topical breadth of follow-up questions—but only when humans retain epistemic authority and ethical oversight. Our core contribution is the first empirical demonstration of how LLMs function as *collaborators*—not substitutes—in qualitative data collection, revealing their impact mechanism on data quality and establishing a methodological foundation and practical pathway for AI-enhanced qualitative research.

Evaluating AI's complementary role to human interviewersExploring human-AI role division in qualitative researchInvestigating AI-generated follow-up questions in interviews

Requirements Elicitation Follow-Up Question Generation

Jul 03, 2025
YS
Yuchen Shen
🏛️ Carnegie Mellon University

During requirements elicitation interviews, experienced interviewers often struggle to formulate high-quality follow-up questions in real time due to domain knowledge gaps, high cognitive load, and information overload. Method: We propose an error-guided prompting framework that integrates a taxonomy of common interviewer errors with the GPT-4o large language model, enabling dynamic generation of contextually relevant questions grounded in受访者 utterances. Structured prompting enhances question relevance and informativeness. Contribution/Results: Empirical evaluation shows that LLM-generated questions match human-authored questions in clarity, relevance, and informativeness—and significantly outperform baseline LLM approaches when guided by error categories. Our approach establishes an interpretable, reusable generative paradigm for intelligent requirements engineering assistance, effectively alleviating interviewer cognitive burden while improving both the quality and efficiency of requirements elicitation.

Addressing challenges like domain unfamiliarity and cognitive load in interviewsEvaluating LLM-generated questions versus human-authored questions for qualityGenerating real-time follow-up questions for requirements elicitation interviews

This study addresses the limited depth of follow-up questioning in semi-structured interviews, which often stems from interviewers’ high cognitive load and constrained domain knowledge. Through a Wizard-of-Oz experiment integrating GPT-4o in a human-in-the-loop setting, the research enables human interview日晚间 to selectively adopt and edit AI-generated probes during real-time conversations, preserving human oversight while exploring viable models for AI-assisted interviewing. The work systematically identifies five interwoven ethical risks inherent in such AI-augmented interactions: harmful language, attentional distraction, unequal participation, blurred accountability, and privacy compliance concerns. Building on these findings, the study proposes design and governance recommendations centered on safety, respect, and accountability, offering an empirical foundation and guiding principles for the development of responsible AI interview systems.

AI ethicsAI-assisted interviewingfollow-up questions

AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers

Sep 16, 2024
AW
Alexander Wuttke
🏛️ LMU Munich | University of Mannheim | University of Oxford

Traditional survey methods face a trade-off between depth and scalability: structured questionnaires scale well but lack expressive flexibility, whereas in-depth interviews yield rich insights yet are labor-intensive and difficult to scale. Method: This study conducts the first controlled experimental evaluation of large language models (LLMs) as adaptive, conversational interviewers—specifically for political topics—comparing AI- and human-administered interviews across data quality, participant engagement, and operational efficiency. We propose a design framework that reconciles standardization with conversational adaptability, integrating structured questionnaire logic, real-time response generation, and multi-dimensional evaluation metrics (e.g., protocol adherence, response quality, engagement). Contribution/Results: AI-conducted interviews achieve data quality comparable to human interviews, demonstrate substantially improved scalability, and elicit positive participant feedback. This work establishes a novel paradigm for high-fidelity, large-scale qualitative data collection in the social sciences.

Assessment of AI Conversational Interviewing performance and improvement opportunitiesPotential of LLMs to replace human interviewers for scalable conversational interviewsTrade-off between depth and scale in traditional survey methods

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

Feb 02, 2024
ZR
Zeeshan Rasheed
🏛️ Tampere University | Jyväskylä University | University of Helsinki | Lancaster University Leipzig | Free University of Bozen Bolzano

Qualitative data analysis in software engineering faces challenges including time intensity, poor reproducibility, and difficulty ensuring inter-rater reliability; the potential of large language models (LLMs) for human–AI collaboration in such tasks remains underexplored. This paper introduces the first explainable multi-agent framework tailored for qualitative research, enabling automated coding, theme extraction, and cross-textual synthesis via role-based task decomposition, prompt engineering, iterative validation, and a closed-loop human feedback mechanism. The architecture preserves human oversight and ensures analytical traceability, overcoming LLM limitations in low-shot, high-reliability settings. Empirical evaluation demonstrates a 3.2× improvement in analysis efficiency, scalability to hundreds of interviews, 89.7% accuracy in theme identification, and strong endorsement by domain experts.

Addressing time-intensive manual analysis that compromises validityAutomating qualitative data analysis using LLM-based multi-agent systemsDeveloping AI-human collaboration for qualitative research automation

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic, quantifiable, and reproducible evaluation methods for assessing the interviewing capabilities of large language models (LLMs) in conversational requirements elicitation. To this end, it introduces the first standardized automated evaluation framework tailored to requirements-gathering interviews, comprising a dataset of 101 scenarios, high-fidelity LLM-driven simulated users, and an automatic task evaluator that achieves strong alignment with real-user interactions and expert judgments across multiple dialogue quality dimensions. Empirical results reveal that current state-of-the-art models uncover fewer than half of implicit requirements, exhibit particularly poor performance on stylistic requirements, and tend to generate effective questions predominantly in the latter stages of dialogues.

conversational AIimplicit requirementsinterview competence

This study addresses the challenges interviewers face in requirements elicitation interviews, where balancing comprehensive topic coverage, active listening, and adaptive follow-up questioning is difficult, compounded by a lack of effective script execution tracking. To overcome these limitations, this work proposes the first end-to-end AI-assisted interview framework, integrating business-goal-driven theoretical script generation, real-time topic coverage monitoring via natural language processing, and an on-demand dynamic probing mechanism. Experimental results demonstrate that the proposed approach significantly improves script quality (92.8 vs. 74.8), probing depth (3.43 vs. 1.15 probes per topic), and granularity of the resulting requirements models (proportion of leaf-level goals: 0.653 vs. 0.598). Notably, 86% of users identified real-time topic tracking as the most practically valuable feature.

AI assistanceinterviewingrequirements elicitation

This study addresses the high labor costs of semi-structured interviews, the opaque behavioral mechanisms of existing multimodal large language models (MLLMs), and the challenges in establishing trust during human–AI interaction. To this end, the authors developed InterviewBot—a system that integrates researcher-designed interview protocols—and conducted the first empirical analysis of MLLM turn-by-turn behaviors in a real-world deployment setting. Leveraging voice-driven real-time MLLM interaction, a semi-structured interview framework, and inductive qualitative methods, the study identifies four data collection failure modes, including information loss and premature termination, and reveals that only 4.9% of AI-generated questions constituted probing follow-ups while 28.7% violated single-question instructions. The findings uncover three key socio-dynamic mechanisms—disclosure calibration, institutional legitimacy, and conversational anchoring—and distill design principles centered on enhanced depth control and non-scripted listening.

breakdownshuman-AI interactionMLLM-led interviews

Existing LLM-based interview systems struggle to balance predefined topic coverage with adaptive exploration, limiting the scalable acquisition of high-quality qualitative user insights. This work proposes a multi-agent LLM architecture that frames adaptive semi-structured interviewing as a utility optimization problem, formally defining interview utility as a trade-off among topic coverage, discovery of novel insights, and conversational cost. The system dynamically plans high-expected-utility questions through simulated dialogue rollouts. Experiments demonstrate that, in LLM simulations, the approach improves topic coverage by 4.7% and yields richer insights in fewer turns. A user study with 70 participants further validates that domain experts recognize the method’s ability to uncover high-quality insights in professional contexts that existing approaches fail to capture.

adaptive interviewingemergent themeslarge language models

Hot Scholars

GN

Graham Neubig

Carnegie Mellon University, All Hands AI
Natural Language ProcessingMachine LearningArtificial Intelligence
DN

David Nguyen

Engineer at Saolasoft Inc.
Data MiningComputer NetworkingBlockchain
EB

Erik Brynjolfsson

Professor at Stanford; NBER; Stanford Digital Economy Lab
EconomicsInformation EconomicsEconomics of AIProductivity
DY

Diyi Yang

Stanford University
Computational Social ScienceNatural Language ProcessingMachine Learning