Score
Designs and builds structured taxonomies and catalogs that classify and organize risks, vulnerabilities, and safety concerns into categories, dimensions, and hierarchies. Specifies category granularity and relationships, maps taxonomy entries to metrics, controls, or mitigations, and adapts the taxonomy to the project or organizational context and intended uses.
This study addresses the challenges of extracting structured risk factors from corporate 10-K filings, particularly inconsistencies with predefined hierarchical taxonomies and the absence of continuous optimization mechanisms. The authors propose an end-to-end, three-stage framework: first, a large language model (LLM) extracts risk factors with source citations; second, semantic embeddings map these factors to a taxonomy; and third, an LLM-as-a-judge mechanism filters erroneous matches. Additionally, an AI agent is introduced to autonomously diagnose and iteratively refine the taxonomy. Experiments on S&P 500 company filings demonstrate a 63% increase in within-industry risk similarity (Cohen’s d = 1.06, AUC = 0.82) and a 104.7% improvement in embedding separation between categories, confirming the method’s effectiveness and generalizability.
This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.
Security logs are inherently unstructured, semantically ambiguous, and ill-suited for deep reasoning. Method: We propose an ontology-driven, large language model (LLM)-augmented knowledge graph construction framework. It integrates a cybersecurity ontology—aligned with NIS 2 and EU classification standards—with semantic log parsing and LLM-enhanced entity-relation extraction to automate log-to-knowledge-graph mapping; ontology constraints improve LLM output accuracy and interpretability. Contribution/Results: This work establishes the first end-to-end semantic–knowledge co-analytical paradigm for security logs. Experiments demonstrate an 18.7% improvement in log structuring F1-score and significantly enhanced threat contextual reasoning. The framework enables cross-system data interoperability and provides a verifiable semantic foundation for intelligent threat detection and regulatory compliance auditing.
Addressing the challenge of multi-label automatic annotation for large-scale, hierarchical classification systems in software requirements engineering, this study proposes a sentence-level zero-shot classification paradigm to circumvent the high annotation costs associated with supervised training. We introduce the first industrial-scale requirements annotation benchmark comprising 769 taxonomy labels and systematically demonstrate a strong negative correlation between the number of taxonomy leaf nodes and classification recall. We further propose a zero-shot multi-label classification method leveraging SBERT sentence embeddings, achieving significant improvements in recall. Empirical evaluation reveals that hierarchical strategies yield no consistent performance gain across settings. Our work validates the effectiveness and feasibility of zero-shot learning for large-scale requirements classification, offering a scalable, low-human-effort automation solution for requirements tracing. (138 words)
This paper addresses the lack of a unified classification framework for the structural properties of relational abstractions. Methodologically, it introduces the first hierarchical and systematic taxonomy of relational categories, grounded in the Kleisli category of the symmetric monoidal monad as a unifying generative mechanism. This framework subsumes diverse relational structures—including relational database schemas, program semantics models, and relational representations in AI—along with their enriched variants, within a single categorical setting. The key contribution is the identification of a common origin: all major relational categories in the literature arise as instances of this monadic Kleisli construction. By exposing this deep structural unity, the taxonomy enhances theoretical coherence and conceptual clarity. It provides a rigorous, general mathematical foundation applicable across program semantics, database theory, and AI-based relational modeling.
This work addresses the lack of a systematic hierarchical taxonomy for GitHub repositories, where existing tag-based mechanisms are flat, inconsistent, and sparsely populated. The authors propose the first end-to-end framework for automatically generating a hierarchical classification of repositories by integrating knowledge from large language models with empirical repository distributions. Their approach employs a multi-agent architecture—comprising designer agents that construct taxonomic dimensions and classifier agents that assign projects—augmented with an iterative self-correction mechanism and a novel hierarchical path evaluation strategy. Evaluated on a benchmark of 2,001 repositories, the method achieves a Taxonomy Quality Factor (TQF) of 83.13%, outperforming the best baseline by 15 percentage points. In downstream tasks, it attains 85.71% precision at rank 1 for alternative discovery, surpassing human-curated lists and substantially improving retrieval efficiency, while also uncovering evolutionary trends in domains such as AI and machine learning.
This study addresses the high cost and expert dependency of manual taxonomy construction in software engineering (SE) by conducting the first systematic, multi-dimensional empirical evaluation of large language model (LLM)-driven automatic classification in this domain. Leveraging two representative approaches—TnT-LLM and CLIMB—and five state-of-the-art LLMs across seven human-annotated SE paper datasets, the work analyzes performance along key dimensions including classification quality, alignment with expert judgments, reliability, and efficiency. Results reveal that TnT-LLM achieves near-human classification quality but incurs high computational cost and structural complexity, whereas CLIMB offers 15–40× faster inference and 8–49× lower cost at the expense of reduced accuracy in tasks requiring deep technical reasoning. The findings elucidate critical trade-offs among quality, cost, and complexity, providing actionable guidance for method selection in practice.
This study addresses the multidimensional risks—operational, security, and governance-related—that enterprises face when deploying large language models, noting that existing open-source tools are fragmented and fail to comprehensively cover authoritative risk taxonomies. To bridge this gap, the work proposes a structured mapping protocol that automatically aligns the capabilities of 21 prominent open-source tools with the 32 subcategories of the MIT AI Risk Framework, leveraging retrieval-augmented generation (RAG) and LLM-based parsing. The protocol’s validity is substantiated through source code and documentation analysis, majority voting, and inter-rater reliability assessment using Fleiss’ Kappa (κ = 0.509, F1 = 75.5%). Findings reveal a pronounced overconcentration of current tools on technical controls, with significant gaps in governance, legal, and market risk domains, thereby providing an empirical foundation for developing layered AI risk mitigation architectures.
This study addresses the absence of a unified and comprehensive classification framework for crypto assets, which hampers informed investment decisions and effective regulatory evaluation. To bridge this gap, the paper proposes a multidimensional taxonomy that integrates technical design, market structure, and regulatory considerations. For the first time, it systematically incorporates dimensions such as technical standards, degrees of resource centralization, asset functionality, legal attributes, and minting and yield mechanisms. Through theoretical derivation, regulatory analysis, and case studies, the framework maps the top 100 mainstream crypto assets, revealing underlying control patterns even in ostensibly decentralized assets. It also accommodates ambiguous edge cases and identifies recurring design paradigms, thereby offering a practical analytical tool for regulatory risk assessment, cross-asset comparison, and the development of digital financial platforms.
This study addresses the critical issue of irreversible classification errors in manual data entry by small and medium-sized enterprises (SMEs), where semantically or morphologically similar categories often distort key performance indicators and mislead decision-making. To mitigate this, the authors propose ISEC, a novel preventive ordinal metric that integrates semantic embeddings, weighted Damerau-Levenshtein edit costs, and empirical category frequencies into a scalable, confusion-aware ranking framework. By leveraging a vector database for efficient similarity search, the method achieves substantial computational gains. Empirical validation across three heterogeneous datasets—legal records, retail inventory, and metalworking catalogs—demonstrates a 195-fold speedup over brute-force computation while maintaining high accuracy, offering SMEs a practical and efficient tool for robust data governance.