🤖 AI Summary
This study addresses the challenge of organizing vast volumes of unstructured, multilingual professional skill statements by proposing a hybrid knowledge graph construction approach that integrates top-down and bottom-up strategies. Anchoring large language models to the Wikidata multilingual knowledge graph and incorporating an agent-based reflection mechanism, the method dynamically aligns known entities, creates nodes for emerging skills, and establishes relationships through a five-stage pipeline. Leveraging techniques such as multilingual normalization, active curation, and deduplication, the system yields a scalable, interpretable, and self-repairing knowledge graph. The resulting resource encompasses a comprehensive skill ontology and structured taxonomy across five European languages, effectively managing noisy textual inputs and continuously recovering unmapped concepts.
📝 Abstract
Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model (LLM) in the Wikidata multilingual Knowledge Graph (KG) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata. Unlike rigid top-down methods or fragmented bottom-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages. Ultimately, this pipeline provides a highly scalable, explicable, and self-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text.