Score
Designs and implements systems that automatically identify and normalize the skills required by a project by parsing project descriptions and requirement texts into structured skill lists using techniques such as information extraction and large language models. Also develops methods to map free-text requirements to standardized skill taxonomies and to reduce manual annotation and curation effort for downstream matching and analytics.
This paper addresses the absence of systematic surveys and methodological frameworks for natural language processing (NLP) applications in human resources (HR). To this end, it introduces the first comprehensive NLP cross-task taxonomy tailored to the full HR lifecycle—spanning eight core tasks: recruitment, person–job matching, employee management, and others. Methodologically, it integrates pretrained language models, information extraction, semantic matching, and multi-task learning, augmented by HR-specific knowledge graphs and rule-based constraints, thereby establishing a novel “person–job alignment”-driven NLP paradigm. The study systematically reviews over 50 publicly available datasets and benchmarks, elucidating how domain-specific challenges—e.g., sparse annotations, contextual ambiguity, and regulatory constraints—drive NLP innovation. It identifies 12 frontier research directions and critical technical bottlenecks. The work delivers both a unifying theoretical framework for academia and an actionable technology roadmap for industry deployment.
This work addresses the inefficiency and insufficient accuracy inherent in extracting and classifying requirements from semi-structured documents within traditional requirements engineering. To overcome these limitations, the authors propose ReXCL, an end-to-end automated tool that integrates heuristic rules with predictive modeling for requirement extraction and employs an encoder-based deep learning architecture with adaptive fine-tuning to achieve high-precision classification. The output of ReXCL is designed for seamless integration into mainstream requirements engineering tools. Empirical evaluation in real-world requirements engineering scenarios demonstrates that the proposed approach significantly enhances both processing efficiency and classification accuracy, confirming its effectiveness and practical utility.
To address the challenges of unstructured requirements documents, low efficiency, and error-proneness in manual processing, this paper proposes a dual-module framework for requirements engineering automation. The framework integrates heuristic text rules with an adaptively fine-tuned Transformer encoder to achieve end-to-end schematization of semi-structured requirements—jointly performing information extraction, standardized schema mapping, and automatic semantic labeling from raw textual input. It supports direct export to mainstream requirements management tools and seamless integration into existing RE toolchains. As the first systematic solution targeting structured transformation of requirements documentation, it achieves an average 18.7% improvement in F1-score across multiple industrial datasets, significantly enhancing classification accuracy and engineering reusability.
This study addresses the scarcity of publicly available course–occupational skill alignment data in educational skill recommendation systems. We present the first fine-grained dataset that aligns graduate-level courses with skills from the European Skills/Competences, Qualifications and Occupations (ESCO) framework—specifically for Systems Analysts and Management and Organizational Analysts. The dataset integrates human-annotated labels and synthetically generated data at both course title and sentence levels, accompanied by a standardized annotation guideline. Leveraging this resource, we train BERT-based language models to perform bidirectional semantic retrieval between courses and skills. Baseline models achieve an F1 score of 87% on the annotated subset, demonstrating the feasibility of the task and filling a critical gap in skill-mapping data on the educational side.
Requirements engineering (RE) education lacks low-cost, high-fidelity environments for interview training. Method: This study proposes an LLM-based interactive virtual client pedagogy, constructing conversational, scalable, LLM-driven virtual stakeholders to enable real-time, online interview practice. It integrates a human–AI collaborative interview framework and hybrid assessment—combining qualitative feedback with quantitative behavioral analysis. Contribution/Results: This is the first systematic application of LLMs to cultivate dynamic interviewing competencies in RE education. The approach maintains technical correctness and efficiency while significantly enhancing student immersion and perceived authenticity. An empirical evaluation shows 87% of students prefer this mode, validating its innovation in scalability, accessibility, and pedagogical effectiveness.
To address challenges in requirements engineering—including difficulty in identifying relationships among natural language requirements, high manual annotation costs, and poor domain adaptability—this paper proposes an NLP-driven, systematic relation extraction framework. It is the first to integrate a requirements relationship ontology with multi-paradigm NLP techniques: dependency parsing, semantic role labeling, named entity recognition, BERT-based supervised fine-tuning, and retrieval-augmented methods. A unified classification-based evaluation framework is established to clarify core challenges and evolutionary pathways. The framework supports major requirement relations (e.g., *refines*, *conflicts*) and enables reusable, extensible relation modeling. Experimental results demonstrate significant improvements in automation capability and accuracy for large-scale adaptive requirements management systems, thereby strengthening requirements evolution analysis and consistency verification.
This work addresses the challenge that large language models struggle to effectively follow textual skill instructions in long-context scenarios, which hinders practical skill deployment. To overcome this limitation, the authors propose ParametricSkills, a novel framework that, for the first time, dynamically converts free-form textual skills into LoRA adapter parameters at test time, enabling context-independent skill invocation and establishing a new paradigm for test-time continual learning. Leveraging a large-scale skill repository and invocation trajectories derived from OpenCode, a hypernetwork is trained to map textual descriptions to parameterized skills. Evaluated across six software engineering subtasks, the method outperforms in-context learning by an average of 6.44 points under DeepSeek-V4-Flash assessment, with significant improvements in both BERTScore and F1 metrics.
Current agent skills are predominantly represented as unstructured text, hindering efficient parsing, reuse, and reasoning, thereby limiting automated skill management. This work proposes a Schedule–Structure–Logic (SSL) tripartite representation framework, inspired by linguistic knowledge representation theories, which explicitly decouples skills into scheduling signals, execution structures, and logical actions. Leveraging large language models, a skill normalizer integrates memory organization packets, script theory, and conceptual dependency theory to automatically generate SSL representations. Experimental results demonstrate that the proposed approach significantly outperforms text-based baselines, improving Mean Reciprocal Rank (MRR) from 0.573 to 0.707 in skill discovery and macro F1-score from 0.744 to 0.787 in risk assessment, thereby substantially enhancing skill retrievability, auditability, and operability.
研究通过两个实证研究评估了大型语言模型在需求工程中的应用,覆盖了从需求分类到可追溯性链接识别等五项活动,旨在解决需求信息提取的难题。
This study addresses the unclear adaptation patterns and impacts of large language model (LLM) agent skills when reused in downstream applications. Through an empirical analysis of 1,126 adaptation instances from six prominent skill repositories, the work systematically characterizes LLM skill adaptation behaviors and constructs a taxonomy comprising 46 patterns grouped into 13 families. The research uncovers critical phenomena including a “reuse paradox,” strong cross-component dependencies, and the introduction of security-sensitive content in nearly one-fifth of adaptations. It further identifies prevalent challenges such as logic rewriting, fixing discoverability issues, and cross-tool or cross-language translation. These findings offer new empirical insights and foundational support for improving skill design, standardizing interfaces, and enabling automated adaptation of LLM-based agents.
This work proposes an automated approach leveraging large language models (LLMs) to generate UML class diagrams from natural language requirements, aiming to reduce manual effort in software design. The method employs chain-of-thought prompting to extract domain entities, attributes, and relationships, followed by structured diagram generation using PlantUML. A novel dual-validation framework is introduced: one component utilizes LLMs such as Grok and Mistral as judges for automated quality assessment, while the other incorporates expert human evaluation. Experimental results demonstrate that the generated class diagrams achieve high performance across five dimensions—including completeness and correctness—and show strong alignment between LLM-based evaluations and expert judgments, thereby confirming the feasibility and reliability of LLMs in both automated modeling and quality assessment.