Score
Designs and develops competency frameworks and models that specify the skills, behaviors, and proficiency levels required for roles, and maps those competencies to job functions and role definitions. Builds competency assessment frameworks and calibration processes to measure, compare, and validate proficiency, including competence mapping and role-specific competency modeling for evaluation and development.
Competency modeling is widely used in human resource management to select, develop, and evaluate talent. However, traditional expert-driven approaches rely heavily on manual analysis of large volumes of interview transcripts, making them costly and prone to randomness, ambiguity, and limited reproducibility. This study proposes a new competency modeling process built on large language models (LLMs). Instead of merely automating isolated steps, we reconstruct the workflow by decomposing expert practices into structured computational components. Specifically, we leverage LLMs to extract behavioral and psychological descriptions from raw textual data and map them to predefined competency libraries through embedding-based similarity. We further introduce a learnable parameter that adaptively integrates different information sources, enabling the model to determine the relative importance of behavioral and psychological signals. To address the long-standing challenge of validation, we develop an offline evaluation procedure that allows systematic model selection without requiring additional large-scale data collection. Empirical results from a real-world implementation in a software outsourcing company demonstrate strong predictive validity, cross-library consistency, and structural robustness. Overall, our framework transforms competency modeling from a largely qualitative and expert-dependent practice into a transparent, data-driven, and evaluable analytical process.
To address poor interoperability, limited adaptability, and insufficient semantic understanding in traditional skill management systems amid rapid labor market transformation, this paper proposes an ontology-based skill management framework. The framework establishes a unified, multi-source skill ontology model formalized in RDF/OWL and leverages semantic reasoning to enable structured modeling and dynamic linking of skills, occupations, and training programs. Its key innovations include cross-domain skill alignment, automated job–competency matching, personalized learning recommendations, and interpretable career pathway planning. Empirical validation across recruitment, vocational education, and lifelong learning scenarios demonstrates significant improvements in matching accuracy and system scalability. The framework provides a reusable semantic infrastructure for skill governance in the digital era.
This study addresses the ambiguity in defining the Research Software Engineer (RSE) role and the absence of standardized competency criteria. Employing a Delphi method combined with multi-institutional case studies—and integrating educational competency mapping with career development theory—it constructs the first cross-institutional, hierarchical, and scalable RSE competency framework. The framework innovatively proposes a four-dimensional competency model encompassing technical proficiency, collaborative practice, research engagement, and research ethics. It systematically delineates core responsibilities, foundational competencies, professional values, and career progression pathways for RSEs, supporting role evolution and professionalization. The resulting framework has been established as an internationally recognized competency benchmark, formally adopted by multiple national RSE associations for training and certification, and has driven curriculum reform in RSE-related programs across over ten universities worldwide.
Current virtual team competency assessment in remote collaboration relies heavily on subjective self-reports and fragmented evaluations, failing to capture authentic behavioral patterns or pinpoint actionable improvement areas. Method: Grounded in group dynamics theory, this study proposes a novel three-dimensional behavioral metric system integrating task-oriented (group/individual) and socio-relational dimensions. Using Critical Incident Technique (CIT), focus group interviews, and configurational analysis, observable and interpretable behavioral markers were empirically derived from engineering students’ collaborative activities. Contribution/Results: The resulting assessment framework is operationally feasible, iteratively refinable, and diagnostically precise—enabling personalized competency development. Validated within engineering education contexts, it demonstrates both empirical effectiveness and contextual applicability for formative, behaviorally anchored virtual team assessment.
Contemporary computer science education suffers from theory-heavy curricula that lack explicit modeling of skill dependencies. Method: This paper proposes a skill-centered dual-tree curriculum structuring framework comprising Skill Trees (explicitly encoding skill hierarchies and dependencies) and Concept Trees (representing foundational conceptual mechanisms underpinning skills). We formally define a computable dual-tree model—overcoming traditional theory-oriented limitations—to enable skill dependency modeling and automated pedagogical path planning. Our methodology includes skill graph construction, concept–skill mapping design, structured curriculum planning, and empirical educational evaluation. Contribution/Results: Applied in an undergraduate database course, the framework significantly reduced students’ cognitive confusion and perceived learning pressure, while decreasing average time-to-mastery for target skills.
This work addresses the lack of systematic evaluation of agent skills' practical utility in cross-domain, multi-model settings. We propose the first scalable evaluation framework specifically designed for agent skills, enabling skill developers to define custom tasks and evaluation dimensions grounded in real-world scenarios. The framework quantifies the enhancement provided by skills to large language model (LLM) agents through metrics of instruction following and task completion. It integrates automated task generation, scoring rules, and comparative experiments across 19 commercial and open-source LLMs, yielding a benchmark comprising 1,000 diverse tasks. Experimental results reveal significant disparities among models in adhering to skill instructions, thereby validating the effectiveness of skills in guiding agent behavior. The benchmark dataset is publicly released to foster further research in this area.
Current evaluation methods struggle to finely assess how large language model agents actually utilize reusable skills and why they fail. This work proposes "skill coverage" as a test adequacy metric tailored to agent skills, translating natural language skill instructions into semi-structured behavioral constraints and evaluating whether these constraints are covered and successfully executed based on execution traces. This approach decouples skill usage from task outcomes, enabling actionable failure attribution. Experiments on SkillsBench reveal that existing agents cover only 38.66%–45.51% of skill constraints; further, reinforcing skills based on failed constraints yields an average task recovery rate of 16.0% across previously failed tasks.
This work addresses the model-agnostic nature of existing skill libraries, which overlooks the substantial capability differences among large language models and consequently leads to highly model-dependent skill effectiveness. To resolve this issue, the authors propose the MASA framework, which achieves adaptive alignment between skills and models through a two-stage process: first, it optimizes skills via hierarchical skill evolution—integrating hill-climbing with UCB tree search—guided by environmental feedback and model capability profiling; second, it trains a lightweight model-conditional rewriter to generalize the optimized skills to new tasks. Notably, MASA requires no modification of model weights, is the first to explicitly identify and mitigate model dependency in skill effectiveness, and enables zero-shot transfer. Experiments across three interactive environments and four mainstream models demonstrate that MASA consistently outperforms strong baselines by up to 25.8 points on average while achieving superior performance over larger teacher models at lower inference cost.