Score
Designs and implements systems and workflows for creating, consolidating, organizing, and maintaining structured repositories of knowledge, covering content ingestion, deduplication, versioning, metadata and ontology management, and lifecycle policies. Builds or analyzes retrieval and indexing mechanisms, search and access controls, and taxonomy or skill-mapping processes to ensure knowledge is accurate, discoverable, and aligned with defined skill schemas.
To address the low efficiency of code understanding and file localization in large-scale software repositories, this paper proposes a graph-augmented hybrid retrieval framework. First, it leverages large language models (LLMs) to extract semantic summaries and generate fine-grained vector embeddings, while integrating static analysis to construct a knowledge graph encoding syntactic–semantic relationships—including inheritance, method calls, and references. Second, it introduces a graph-aware retrieval expansion mechanism that jointly optimizes semantic similarity matching and subgraph traversal, supporting LLM-generated constrained natural language queries and interpretable reasoning. Experimental results demonstrate significant improvements in accuracy and robustness for problem-driven file retrieval across multiple open-source projects. The approach achieves a substantial leap in automated code understanding capability and establishes a scalable, knowledge-infused infrastructure for intelligent development toolchains.
To address poor interoperability, limited adaptability, and insufficient semantic understanding in traditional skill management systems amid rapid labor market transformation, this paper proposes an ontology-based skill management framework. The framework establishes a unified, multi-source skill ontology model formalized in RDF/OWL and leverages semantic reasoning to enable structured modeling and dynamic linking of skills, occupations, and training programs. Its key innovations include cross-domain skill alignment, automated job–competency matching, personalized learning recommendations, and interpretable career pathway planning. Empirical validation across recruitment, vocational education, and lifelong learning scenarios demonstrates significant improvements in matching accuracy and system scalability. The framework provides a reusable semantic infrastructure for skill governance in the digital era.
In the era of large language models, traditional record-centric data engineering struggles to meet the demand for organizational knowledge as executable infrastructure. This work proposes a novel paradigm—knowledge architecture—that systematically reimagines core data engineering mechanisms by upgrading ETL, data lineage, and catalogs into knowledge ingestion, change detection, provenance, and knowledge catalogs. It introduces knowledge views and a three-tier layered model (raw–refined–operational) to structure knowledge effectively. By integrating emerging standards such as LLM Wiki and Open Knowledge Format (OKF), this study formally defines knowledge architecture for the first time and establishes a theoretical framework that supports knowledge representation, governance, and operational delivery, enabling direct invocation of organizational knowledge by humans, agents, workflows, and models alike.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
Knowledge Organization Systems (KOS) in academia exhibit high heterogeneity in scope, structure, quality, and interoperability, impeding effective research information organization and utilization. To address this, we propose the first five-dimensional evaluation framework—covering scope, structure, maintenance, usage, and interoperability—for 45 representative KOS, including glossaries, thesauri, taxonomies, and ontologies. Integrating qualitative expert interviews with structured metadata analysis, our study systematically characterizes cross-disciplinary heterogeneity. Results reveal significant disparities across KOS in scale, quality, and interoperability, identifying three core challenges: lagging standardization, insufficient dynamic evolution mechanisms, and difficulties in cross-domain alignment. Based on these findings, we articulate a novel integrative paradigm for research knowledge representation. This work provides both theoretical foundations and practical guidelines for AI-driven scholarly knowledge management.
Current agent capabilities lack large-scale infrastructure for production, governance, and evolution akin to Wikipedia and GitHub. This work proposes SkillWiki, the first dynamic knowledge infrastructure that enables symbiotic co-evolution of knowledge, skills, and execution experiences. By integrating knowledge extraction, skill assetization, provenance tracking, and an execution-driven feedback loop, SkillWiki establishes an end-to-end framework for the full lifecycle management of skills, supporting reusable organization, verifiable provenance, and continuous evolution. The system has been fully validated through a pipeline spanning knowledge ingestion, skill generation, and execution-driven refinement. Both the codebase and the system are publicly open-sourced.
This study addresses the lack of systematic understanding regarding the creation, reuse, customization, and maintenance of AI agent skills as reusable software artifacts. Treating AI skills as engineered software artifacts for the first time, the authors conduct an empirical investigation based on over 40,000 skill instances drawn from public registries and GitHub repositories, employing a mixed-methods approach that combines large-scale data mining, LLM-driven SWEBOK-based knowledge classification, topic modeling, and qualitative coding. The findings reveal that 53% of reused skills remain unmodified, with reuse predominantly involving one-time copying; customization primarily serves to adapt skills to local environments; and evolution tends to follow an incremental addition pattern while preserving highly stable behavioral contracts. The study identifies six content categories of skills and six modification themes, establishing an empirical foundation for the engineering-oriented management of AI skills.
This study addresses the challenge of organizing vast volumes of unstructured, multilingual professional skill statements by proposing a hybrid knowledge graph construction approach that integrates top-down and bottom-up strategies. Anchoring large language models to the Wikidata multilingual knowledge graph and incorporating an agent-based reflection mechanism, the method dynamically aligns known entities, creates nodes for emerging skills, and establishes relationships through a five-stage pipeline. Leveraging techniques such as multilingual normalization, active curation, and deduplication, the system yields a scalable, interpretable, and self-repairing knowledge graph. The resulting resource encompasses a comprehensive skill ontology and structured taxonomy across five European languages, effectively managing noisy textual inputs and continuously recovering unmapped concepts.
Current large language model (LLM) workflow systems lack semantic representation and persistence mechanisms for workflows themselves, hindering inspectability, recoverability, and auditability. This work proposes a language-agnostic conceptual model inspired by Lisp that treats workflows as knowledge objects rather than mere execution traces. By leveraging symbolic forms, object identity, and the notion of live mirrors, the model distinguishes deterministic computation (derive) from LLM-mediated judgment (infer). It further integrates contextual snapshots and capability policies to govern reasoning processes. This approach establishes a semantic persistence framework for LLM workflows, laying preliminary formal foundations and substantially enhancing their inspectability, recoverability, and auditability.