project skill extraction

Designs and implements systems that automatically identify and normalize the skills required by a project by parsing project descriptions and requirement texts into structured skill lists using techniques such as information extraction and large language models. Also develops methods to map free-text requirements to standardized skill taxonomies and to reduce manual annotation and curation effort for downstream matching and analytics.

projectskillextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiency and insufficient accuracy inherent in extracting and classifying requirements from semi-structured documents within traditional requirements engineering. To overcome these limitations, the authors propose ReXCL, an end-to-end automated tool that integrates heuristic rules with predictive modeling for requirement extraction and employs an encoder-based deep learning architecture with adaptive fine-tuning to achieve high-precision classification. The output of ReXCL is designed for seamless integration into mainstream requirements engineering tools. Empirical evaluation in real-world requirements engineering scenarios demonstrates that the proposed approach significantly enhances both processing efficiency and classification accuracy, confirming its effectiveness and practical utility.

requirement classificationrequirement extractionrequirements engineering

ReXCL: A Tool for Requirement Document Extraction and Classification

Apr 10, 2025
PB
Paheli Bhattacharya
🏛️ Bosch

To address the challenges of unstructured requirements documents, low efficiency, and error-proneness in manual processing, this paper proposes a dual-module framework for requirements engineering automation. The framework integrates heuristic text rules with an adaptively fine-tuned Transformer encoder to achieve end-to-end schematization of semi-structured requirements—jointly performing information extraction, standardized schema mapping, and automatic semantic labeling from raw textual input. It supports direct export to mainstream requirements management tools and seamless integration into existing RE toolchains. As the first systematic solution targeting structured transformation of requirements documentation, it achieves an average 18.7% improvement in F1-score across multiple industrial datasets, significantly enhancing classification accuracy and engineering reusability.

Automates extraction and classification in requirement engineeringImproves efficiency and accuracy in managing software requirementsProcesses raw documents into predefined schema using AI

This study addresses the scarcity of publicly available course–occupational skill alignment data in educational skill recommendation systems. We present the first fine-grained dataset that aligns graduate-level courses with skills from the European Skills/Competences, Qualifications and Occupations (ESCO) framework—specifically for Systems Analysts and Management and Organizational Analysts. The dataset integrates human-annotated labels and synthetically generated data at both course title and sentence levels, accompanied by a standardized annotation guideline. Leveraging this resource, we train BERT-based language models to perform bidirectional semantic retrieval between courses and skills. Baseline models achieve an F1 score of 87% on the annotated subset, demonstrating the feasibility of the task and filling a critical gap in skill-mapping data on the educational side.

course-skill matchingdataset scarcityprofessional competencies

Using Large Language Models to Develop Requirements Elicitation Skills

Mar 10, 2025
NL
Nelson Lojo
🏛️ Univ. of California | Univ. of Sevilla

Requirements engineering (RE) education lacks low-cost, high-fidelity environments for interview training. Method: This study proposes an LLM-based interactive virtual client pedagogy, constructing conversational, scalable, LLM-driven virtual stakeholders to enable real-time, online interview practice. It integrates a human–AI collaborative interview framework and hybrid assessment—combining qualitative feedback with quantitative behavioral analysis. Contribution/Results: This is the first systematic application of LLMs to cultivate dynamic interviewing competencies in RE education. The approach maintains technical correctness and efficiency while significantly enhancing student immersion and perceived authenticity. An empirical evaluation shows 87% of students prefer this mode, validating its innovation in scalability, accessibility, and pedagogical effectiveness.

High cost of teaching Requirements Elicitation skills effectively.Proposing LLM-based interactive interviews as a scalable alternative.Traditional methods fail to develop interview skills adequately.

Automated Requirements Relation Extraction

Jan 22, 2024
QM
Quim Motger
🏛️ Universitat Polit`ecnica de Catalunya

To address challenges in requirements engineering—including difficulty in identifying relationships among natural language requirements, high manual annotation costs, and poor domain adaptability—this paper proposes an NLP-driven, systematic relation extraction framework. It is the first to integrate a requirements relationship ontology with multi-paradigm NLP techniques: dependency parsing, semantic role labeling, named entity recognition, BERT-based supervised fine-tuning, and retrieval-augmented methods. A unified classification-based evaluation framework is established to clarify core challenges and evolutionary pathways. The framework supports major requirement relations (e.g., *refines*, *conflicts*) and enables reusable, extensible relation modeling. Experimental results demonstrate significant improvements in automation capability and accuracy for large-scale adaptive requirements management systems, thereby strengthening requirements evolution analysis and consistency verification.

Addressing ambiguity and effort in requirements engineeringAutomated extraction of relations between textual requirementsExploring NLP techniques for efficient relation identification

Latest Papers

What's happening recently
View more

Parametric Skills

Jun 29, 2026

This work addresses the challenge that large language models struggle to effectively follow textual skill instructions in long-context scenarios, which hinders practical skill deployment. To overcome this limitation, the authors propose ParametricSkills, a novel framework that, for the first time, dynamically converts free-form textual skills into LoRA adapter parameters at test time, enabling context-independent skill invocation and establishing a new paradigm for test-time continual learning. Leveraging a large-scale skill repository and invocation trajectories derived from OpenCode, a hypernetwork is trained to map textual descriptions to parameterized skills. Evaluated across six software engineering subtasks, the method outperforms in-context learning by an average of 6.44 points under DeepSeek-V4-Flash assessment, with significant improvements in both BERTScore and F1 metrics.

instruction followinglarge language modelslong-context understanding

Current agent skills are predominantly represented as unstructured text, hindering efficient parsing, reuse, and reasoning, thereby limiting automated skill management. This work proposes a Schedule–Structure–Logic (SSL) tripartite representation framework, inspired by linguistic knowledge representation theories, which explicitly decouples skills into scheduling signals, execution structures, and logical actions. Leveraging large language models, a skill normalizer integrates memory organization packets, script theory, and conceptual dependency theory to automatically generate SSL representations. Experimental results demonstrate that the proposed approach significantly outperforms text-based baselines, improving Mean Reciprocal Rank (MRR) from 0.573 to 0.707 in skill discovery and macro F1-score from 0.744 to 0.787 in risk assessment, thereby substantially enhancing skill retrievability, auditability, and operability.

agent skillsnatural language understandingskill management

This study addresses the unclear adaptation patterns and impacts of large language model (LLM) agent skills when reused in downstream applications. Through an empirical analysis of 1,126 adaptation instances from six prominent skill repositories, the work systematically characterizes LLM skill adaptation behaviors and constructs a taxonomy comprising 46 patterns grouped into 13 families. The research uncovers critical phenomena including a “reuse paradox,” strong cross-component dependencies, and the introduction of security-sensitive content in nearly one-fifth of adaptations. It further identifies prevalent challenges such as logic rewriting, fixing discoverability issues, and cross-tool or cross-language translation. These findings offer new empirical insights and foundational support for improving skill design, standardizing interfaces, and enabling automated adaptation of LLM-based agents.

downstream modificationempirical studylarge language model agents

This work proposes an automated approach leveraging large language models (LLMs) to generate UML class diagrams from natural language requirements, aiming to reduce manual effort in software design. The method employs chain-of-thought prompting to extract domain entities, attributes, and relationships, followed by structured diagram generation using PlantUML. A novel dual-validation framework is introduced: one component utilizes LLMs such as Grok and Mistral as judges for automated quality assessment, while the other incorporates expert human evaluation. Experimental results demonstrate that the generated class diagrams achieve high performance across five dimensions—including completeness and correctness—and show strong alignment between LLM-based evaluations and expert judgments, thereby confirming the feasibility and reliability of LLMs in both automated modeling and quality assessment.

Class Diagram GenerationLarge Language ModelsNatural Language Requirements

Hot Scholars