Score
Designs and builds survey and interview instruments and protocols, including structured questionnaires, semi‑structured interview guides, survey experiment framings, and associated sampling and consent procedures. Conducts and manages administration (e.g., online deployment and interview execution), implements randomization and data-collection workflows, and prepares resulting self-report and behavioral data for analysis and interpretation.
This study addresses the challenges of traditional semi-structured interviews in empirical software engineering, which are often resource-intensive and hindered by cross-time-zone coordination and multilingual barriers. The authors propose a self-administered AI interview approach based on a customized MyGPT model, enabling participants to complete unmoderated interviews via voice in their preferred language, with the system automatically generating structured summaries according to a predefined protocol. As the first work to demonstrate the feasibility of AI-conducted, short-duration, low-risk interviews in this domain, the evaluation shows that 92.4% of 66 submissions met formatting requirements; 90.9% of participants reported a positive experience, 95.5% found the questions clear, and 89.4% expressed willingness to participate again, indicating high acceptability and effectiveness. The study also identifies limitations concerning interview depth and privacy concerns.
This study investigates how question wording in occupational surveys affects the accuracy and linguistic variability of automated job classification. Using a multi-wave German survey experiment, we compare the coding performance of tools including CASCOT and OccuCoDe under two question formats: “job title” versus “occupational tasks,” while quantifying response-level linguistic diversity. Our key contribution is the first empirical demonstration that question format critically moderates automated coding outcomes: the “job title” formulation significantly improves both coding efficiency and accuracy, whereas adding task examples—though enriching response detail—induces lexical homogenization, reducing coding accuracy by 12–18%. These findings uncover a fundamental design-driven mechanism influencing automated occupational classification, establishing that questionnaire structure directly shapes the quality of coded occupational data. The results provide rigorous empirical evidence and methodological guidance for optimizing occupational data collection and coding practices in large-scale social and labor market research.
This study investigates how respondents’ prior survey experiences in online probability panels influence subsequent response behavior, aiming to reduce nonresponse and panel attrition. Using longitudinal panel data, it is the first to systematically apply discrete-time survival analysis to nonresponse modeling—effectively accommodating unbalanced panel structures—while integrating dynamic effects of multi-wave survey experiences (e.g., duration, perceived enjoyment, mode—telephone vs. web—and inter-wave interval) and stable individual traits (e.g., conscientiousness, openness). Results indicate that longer survey duration, lower enjoyment ratings, telephone administration, and extended inter-wave intervals significantly increase refusal risk; conversely, personality traits exhibit robust cross-wave predictive power for response propensity. The findings advance theoretical understanding of panel engagement dynamics and provide empirically grounded, actionable insights for optimizing panel management strategies and enhancing data quality in longitudinal survey research.
Large language models (LLMs) generate personality data lacking the intrinsic heterogeneity observed in human populations, undermining ecological validity in empirical social science research. To address this, we propose the Theory-Driven Personality Structured Interview (PSI) framework, which anchors LLM generation to 357 real human interview transcripts. PSI integrates psychometrically grounded interview design, fine-grained prompt engineering, and behavioral prediction modeling to constrain and guide LLM outputs toward authentic personality diversity. Our key contribution is the first use of theory-anchored structured interviews both as generative constraints and as validation benchmarks, alongside a scalable interview paradigm and systematic evaluation protocol. Experiments demonstrate that PSI significantly enhances the predictive validity of LLM-generated responses on organizational citizenship behavior and counterproductive work behavior tasks—achieving performance statistically indistinguishable from that of human-derived data.
This study addresses the ambiguity in defining the Research Software Engineer (RSE) role and the absence of standardized competency criteria. Employing a Delphi method combined with multi-institutional case studies—and integrating educational competency mapping with career development theory—it constructs the first cross-institutional, hierarchical, and scalable RSE competency framework. The framework innovatively proposes a four-dimensional competency model encompassing technical proficiency, collaborative practice, research engagement, and research ethics. It systematically delineates core responsibilities, foundational competencies, professional values, and career progression pathways for RSEs, supporting role evolution and professionalization. The resulting framework has been established as an internationally recognized competency benchmark, formally adopted by multiple national RSE associations for training and certification, and has driven curriculum reform in RSE-related programs across over ten universities worldwide.
This study addresses the lack of standardized reporting criteria and operational definitions for sample size adequacy and data saturation in software engineering interview research. Through a systematic analysis of 138 interview-based studies published in major software engineering conferences between 2016 and 2025, the authors employ bibliometric and qualitative content analysis methods to code and synthesize reported sample sizes, discussions of saturation, and justifications for methodological choices. This work presents the first large-scale empirical mapping of operational practices regarding interview sample sizes in the field, revealing that most studies recruit between 13 and 24 participants, while samples smaller than 12 are typically confined to specific industrial contexts. Notably, the majority of studies fail to explicitly articulate their criteria for determining data saturation. These findings contribute to enhancing transparency and methodological rigor in qualitative software engineering research.
This study addresses the challenges interviewers face in requirements elicitation interviews, where balancing comprehensive topic coverage, active listening, and adaptive follow-up questioning is difficult, compounded by a lack of effective script execution tracking. To overcome these limitations, this work proposes the first end-to-end AI-assisted interview framework, integrating business-goal-driven theoretical script generation, real-time topic coverage monitoring via natural language processing, and an on-demand dynamic probing mechanism. Experimental results demonstrate that the proposed approach significantly improves script quality (92.8 vs. 74.8), probing depth (3.43 vs. 1.15 probes per topic), and granularity of the resulting requirements models (proportion of leaf-level goals: 0.653 vs. 0.598). Notably, 86% of users identified real-time topic tracking as the most practically valuable feature.
This study evaluates whether large language models (LLMs) can automatically generate questionnaires capable of effectively measuring social attitudes and rivaling established expert-designed scales. Through a within-subjects experimental design, the performance of GPT-4–generated questionnaires—elicited via structured prompts—was systematically compared against validated human-crafted scales across three domains: climate change, immigration, and diversity and inclusion. This work presents the first multi-domain, within-participant comparison between LLM-generated instruments and standard psychometric scales. Results indicate that LLM-generated questionnaires reliably capture major attitudinal divides and are suitable for exploratory, large-scale attitude assessment. However, they exhibit lower resolution in uncovering belief structures and reduced precision in differentiating subpopulations compared to expert-developed scales, suggesting promising yet supplementary utility in social science research.
为解决软件需求收集中的效率与质量难题,提出LadderTeam框架,利用双代理大语言模型自动化用户体验线框图访谈,有效提高反馈的详细度和可执行性。
本文通过回顾2005-2026年间软件工程面试过程,提出基于工作职责的Ticket-Based Pragmatic Assessment Framework,旨在使面试更贴近实际工作。