Score
Designs and implements end-to-end qualitative coding systems by developing codebooks and annotation schemes, writing clear annotation instructions, defining and operationalizing codes and themes, and conducting open- and thematic-coding workflows on qualitative data. Runs and documents manual coding processes including coder training and calibration, qualitative code review, measurement and improvement of interrater agreement, and synthesis of coded results into high-confidence themes and actionable recommendations.
Qualitative evaluation of inductive coding faces significant challenges: conventional metrics are ill-suited for exploratory processes; manual assessment is labor-intensive; and expert-annotated “ground truth” introduces methodological limitations. This paper introduces the first quantifiable evaluation framework for open coding in grounded theory and thematic analysis. Moving beyond the conventional human–AI alignment paradigm, it innovatively integrates stability assessment (Cohen’s kappa, Jaccard similarity) with cross-human–machine comparison as a dual-validation mechanism. Crucially, the framework operates without presupposing ground-truth labels, enabling bias detection and coding quality measurement in human–AI collaborative workflows. Evaluated on two HCI datasets, the framework demonstrates high inter-coder agreement (κ > 0.75), strong output stability (Jaccard similarity > 0.89 across repeated runs), and yields a reusable, AI-augmented coding workflow.
Qualitative analysis in informal settings—such as meeting summaries or personal ideation—lacks rapid, structured computational support. Method: This paper introduces MindCoder, the first lightweight, LLM-driven inductive analysis tool integrating a “code-to-theory” paradigm. Built on GPT-4o, it unifies prompt engineering, iterative human-AI collaboration, and an automated coding pipeline to perform end-to-end open coding, axial coding, and concept generation—balancing analytical efficiency with theoretical rigor. Contribution/Results: A user study (N=12) demonstrates that MindCoder significantly improves analytical flexibility and structural coherence over ChatGPT and Atlas.ti Web AI, reduces task completion time by ~60%, and accelerates consensus formation. Its core innovation lies in embedding grounded theory logic directly into the LLM workflow, enabling interpretable, traceable, and fully automated theory generation without requiring manual coding manuals.
This study addresses the fragmented understanding of “vibe coding” by systematically synthesizing knowledge dispersed across academic and practitioner literature. Employing a unified protocol and a multi-voiced literature review methodology, it analyzes 47 peer-reviewed and gray literature sources from 2022 to 2025, revealing that vibe coding is fundamentally an intent-driven iterative cycle of generation, evaluation, and refinement. The findings indicate that 45% of the reviewed works report short-term productivity gains, with the strongest empirical support for its efficacy in prototyping and UI development. However, significant evidence gaps persist regarding its applicability in production-grade, data-intensive, and safety-critical contexts. This work provides a structured empirical foundation for understanding the evolving role of developers and delineating the appropriate boundaries for vibe coding adoption.
Existing evaluation methods for LLM-generated code comments rely on small-scale datasets and inadequate IR metrics (e.g., BLEU), failing to capture semantic fidelity. Method: We systematically assess GPT-3.5’s Javadoc generation for 23,850 Java code snippets, employing a dual-dimensional evaluation combining quantitative BLEU scoring with qualitative expert human assessment. Contribution/Results: Our study reveals a critical flaw in BLEU: high scores frequently correlate with low-quality, verbatim descriptions, while high-fidelity semantic paraphrasing is systematically penalized. We find that 69.7% of generated Javadocs are semantically equivalent to—or can be refined to match—the original quality, and 22.4% significantly surpass the originals. These results demonstrate that automated metrics alone are unreliable for assessing documentation quality. We advocate human evaluation as the gold standard, with BLEU serving only as a supplementary heuristic—establishing a new, more rigorous paradigm for evaluating code documentation generation.
This study addresses the lack of empirical evaluation regarding whether existing dataset documentation frameworks effectively foster developer reflectivity. Combining mixed-methods thematic analysis with corpus-assisted discourse analysis, the research systematically examines how prevailing documentation frameworks—and their real-world instantiations—cover core dimensions of reflectivity. The findings reveal, for the first time, that current frameworks consistently overlook critical reflective themes. Building on this insight, the authors develop a reflectivity-oriented coding manual and propose an enhanced datasheet template incorporating targeted prompts to elicit deeper reflection. This work offers actionable strategies and practical tools to strengthen the reflective capacity of dataset documentation practices.
This work addresses the persistent challenge clinicians face in translating real-world workflow needs into functional digital health tools due to limited technical expertise and inadequate commercial software support. To bridge this gap, the authors propose “vibe coding”—a method that leverages natural language prompts to guide large language models in collaborative development, thereby lowering technical barriers and enabling non-specialist developers to rapidly prototype solutions tailored to clinical contexts. Integrating clinical workflow analysis, human-AI collaborative programming, and prompt engineering, the study offers practical guidelines, illustrative case studies, and deployment recommendations specifically designed for frontline healthcare professionals. The approach demonstrates both feasibility and practical utility in aligning clinical insights with effective technical implementation.
Existing code documentation often suffers from incompleteness, obsolescence, or inaccuracies, hindering developer comprehension and limiting the performance of large language models (LLMs) in software engineering tasks. This work introduces, for the first time, the concept of “code-document equivalence” and presents Documentary, a novel approach that synthesizes code semantic analysis with LLM-driven natural language generation to automatically produce high-quality, semantically consistent documentation. Experimental results demonstrate that Documentary generates equivalent documentation for 53.4% of functions, significantly improving LLM accuracy on code understanding and editing tasks by 12.8–24.5% compared to both human-written and baseline documentation. Furthermore, developers consistently rate Documentary-generated documentation higher in quality and usefulness.
This study addresses the lack of systematic synthesis in soft skills research within agile software development over the past 25 years, a gap that has hindered the integration of technical and human-centric factors. Through a systematic literature mapping of 97 studies sourced from multiple databases spanning 2000 to 2025, this work constructs the first evolutionary map of soft skills in agile contexts, with a focus on mainstream frameworks such as Scrum. The analysis identifies communication, adaptability, teamwork, and leadership as core soft skills, elucidates their relationships with specific agile roles and methodologies, and highlights a critical research gap concerning role-specific soft skill requirements. These findings offer empirical grounding and theoretical support for advancing agile education, training programs, and organizational practices.