Score
Designs and implements instruments and processes to evaluate individuals’ technical abilities—such as tests, practical tasks, rubrics, scoring algorithms, and data collection protocols—and builds systems to administer them. Analyzes assessment results to quantify competence, estimate reliability and validity, and identify skill gaps for improvement or decision-making.
Contemporary technical hiring practices suffer from stress-induced bias and evidentiary gaps, resulting in distorted competency assessments and compromised fairness. Method: This paper proposes an evidence-driven paradigm for software engineer competency evaluation. It systematically identifies and bridges evidentiary gaps in technical hiring through (1) multi-source behavioral data integration, (2) low-stress, authentic task design, and (3) a verifiable fairness framework grounded in educational measurement, human-computer interaction evaluation, algorithmic fairness auditing, and structured competency modeling. Contribution/Results: The approach yields a scalable, empirically validated hiring effectiveness metric suite. Empirical evaluation demonstrates significant improvements in employer hiring accuracy. Crucially, it establishes a reproducible, auditable foundation for equitable assessment—enhancing both validity and procedural fairness for candidates while enabling rigorous, transparent evaluation of hiring systems.
This study addresses the prevalent challenge of inadequate technical interview preparation among software engineering job seekers. Employing a mixed-methods approach, we conducted a survey with 131 candidates and integrated qualitative and quantitative analyses to examine their preparation behaviors, gaps in educational support, and underlying mechanisms. We empirically identify a critical “education–interview competence gap”: university curricula largely omit authentic interview-oriented training in coding articulation, technical communication, and stress resilience—leading candidates to rely on inefficient self-study, thereby exacerbating anxiety and performance–expectation mismatches. Our key contributions are: (1) the first systematic empirical documentation of this pedagogical–industrial misalignment; and (2) a tripartite collaborative framework—engaging academia, industry, and candidates—to embed interview literacy into computing education and reform assessment toward authenticity. Findings provide data-driven guidance for enhancing computer science pedagogy and recruitment practices.
This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.
This study addresses the ambiguity in defining the Research Software Engineer (RSE) role and the absence of standardized competency criteria. Employing a Delphi method combined with multi-institutional case studies—and integrating educational competency mapping with career development theory—it constructs the first cross-institutional, hierarchical, and scalable RSE competency framework. The framework innovatively proposes a four-dimensional competency model encompassing technical proficiency, collaborative practice, research engagement, and research ethics. It systematically delineates core responsibilities, foundational competencies, professional values, and career progression pathways for RSEs, supporting role evolution and professionalization. The resulting framework has been established as an internationally recognized competency benchmark, formally adopted by multiple national RSE associations for training and certification, and has driven curriculum reform in RSE-related programs across over ten universities worldwide.
This study investigates the impact of generative artificial intelligence (GenAI) on the core competencies required of entry-level software engineers and the corresponding assessment strategies. Through a mixed-methods survey involving 56 educators and 24 hiring professionals, it offers the first systematic comparison of perspectives between academia and industry regarding GenAI usage norms, academic integrity policies, and interview practices. Findings reveal consensus on the importance of critically evaluating AI outputs, responsible GenAI use, and self-directed learning; however, industry respondents consistently perceive recent graduates as lacking foundational programming proficiency, and note that GenAI-integrated evaluation in interviews remains nascent. The study proposes a paradigm shift—from prohibiting GenAI use toward designing higher-order assessment tasks resilient to GenAI assistance—providing empirical grounding for reforming talent evaluation frameworks.
This work addresses the limitations of existing evaluations for large language model (LLM) agents, which predominantly rely on fixed task sets and fail to capture the true value and risks of agent capabilities. The authors propose SkillAudit, a novel skill-centric evaluation framework that decomposes agent skills into skill packages and generates capability-aligned tasks for assessment. Executed within an isolated sandbox environment and judged automatically by LLMs, SkillAudit produces auditable evaluation reports across three dimensions: utility, efficiency/cost, and safety. Its key innovations include a baseline-comparison principle to quantitatively measure utility and cost, and a two-stage safety verification mechanism combining static semantic analysis with dynamic runtime validation. Applied to real-world skill packages across 23 occupational domains, the framework identified security risks in over 7% of skills, demonstrating its effectiveness and necessity.
This study addresses the challenge of delayed and incomplete evaluation in team tabletop exercises (TTX), which often arises due to the open-ended and complex nature of such tasks, hindering effective assessment of team learning outcomes. To overcome this limitation, the authors propose a novel approach that integrates clustering algorithms with large language models (GPT-4o and GPT-5.2) to enable automated, scalable evaluation of team performance. Leveraging action logs and communication transcripts from 81 multinational participants alongside standardized scoring rubrics, the method demonstrates that clustering is computationally efficient and reliable, while GPT-5.2 significantly outperforms GPT-4o in evaluating team communication with lower error rates. All data, tools, and the complete TTX scenario have been open-sourced and integrated into the INJECT platform to support educational applications.