Score
Designs and builds interpretable concept taxonomies and labeled intent catalogs by using large language model outputs to generate, distill, and standardize concise, human‑readable concept names and hierarchical or relational structures. This competence includes creating label‑to‑embedding or label‑to‑session mappings, extracting intent/extent/relational attributes and implications, enforcing consistency across related concepts, and producing guidelines and artifacts for human review and downstream use.
This work addresses the interpretability of large language model (LLM) components—including neurons, attention heads—and abstract features extracted by sparse autoencoders. Method: We propose the first systematic survey framework that unifies generation, evaluation, and dataset curation around natural-language concept descriptions; we introduce a generative approach to produce open-vocabulary descriptions and integrate automated metrics with human evaluation for multi-dimensional validation. Contribution/Results: Our analysis exposes critical deficiencies in existing methods—particularly concerning causal attribution and evaluation rigor—and issues the first explicit call for a causally grounded interpretability paradigm. The study delivers a structured technical roadmap for LLM mechanistic transparency, advancing explainable AI from phenomenological description toward principled, mechanism-based verification.
In domain ontology construction, mapping multi-source terminology to foundational concepts faces three key challenges: high cost and subjectivity of manual approaches, shallow semantic modeling and poor cross-domain consistency of automated methods, and weak interpretability. To address these, this paper proposes the first LLM-driven framework integrating expert calibration with iterative prompt optimization. The framework combines expert-guided annotation, multi-stage prompt engineering, and a human-in-the-loop validation cycle to generate concept links with high confidence and full interpretability. Evaluated on the concept necessity mapping task, it achieves an F1-score of 0.97—substantially surpassing the human baseline (0.68)—and marks the first instance of scalable ontology alignment that simultaneously attains expert-level accuracy and transparent, auditable reasoning.
This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.
This work addresses the limitation that conceptual knowledge in large language models (LLMs) is implicitly encoded through statistical correlations, lacking explicit, structured, and composable representations—hindering model stability, controllability, and alignment with human cognition. To tackle this, the paper introduces the first design-space taxonomy for conceptual structures in LLMs, proposing a “four-stage–two-source” framework that spans training, architecture, inference, and interpretation, combined with internal derivation and external anchoring. The framework systematically integrates probing, dictionary learning, and external knowledge-guided approaches, revealing critical gaps in current research—particularly insufficient exploration of the inference stage, fragmentation across stages, and inconsistent terminology. By shifting the paradigm from passive discovery to active design of conceptual representations, this study provides a theoretical foundation and strategic direction for developing LLMs that are more controllable, interpretable, and cognitively aligned.
Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.
The exponential growth of academic literature poses significant challenges for efficiently constructing comparative tables in survey papers. Existing schema generation methods suffer from ambiguous evaluation criteria and limited editability. To address these issues, this paper proposes an intent-aware schema generation and editing framework: (1) it introduces intent modeling to mitigate semantic ambiguity in comparative dimension identification; (2) it designs an editable generation pipeline enabling on-demand customization of comparison dimensions; (3) it constructs the first benchmark dataset tailored for conditional schema generation; and (4) it integrates LLM-based prompt engineering with lightweight fine-tuning, combining one-shot generation and multi-stage editing strategies. Experimental results demonstrate that intent enhancement substantially improves schema reconstruction accuracy, while the editing mechanism further refines output quality. Notably, our lightweight fine-tuned model achieves performance competitive with state-of-the-art prompting-based large language models.
Current evaluations of large language models predominantly focus on isolated tasks, lacking a systematic capability framework, which hinders cross-study comparability and leaves critical coverage gaps. This work proposes the first three-tier capability taxonomy grounded in human cognitive science, encompassing 14 capability domains and 91 sub-skills, thereby shifting the unit of analysis from tasks to structured competencies. Leveraging multi-model collaborative annotation, consensus mechanisms, and arbitration protocols, the study systematically maps and statistically analyzes nearly 16,000 papers published between 2023 and 2025. The findings reveal a pronounced concentration of research on language and reasoning, with over 60% of capability domains accounting for less than 2% of studies. Additionally, the analysis identifies recurrent co-occurring capability clusters and quantifies their association strengths, offering empirically testable hypotheses for model evaluation, training, and transfer.
This study addresses the challenge that concepts generated by Formal Concept Analysis (FCA) and Relational Concept Analysis (RCA) often rely on technical labels lacking human-interpretable semantic names, thereby hindering domain experts’ understanding and reuse of knowledge. To overcome this limitation, the work proposes a configurable framework that models concept naming as a controlled variability problem, integrating variability modeling with large language models (LLMs). By explicitly governing the exposure of multi-source semantic cues—such as intent, extent, inheritance, and neighboring concepts—the framework generates interpretable names that simultaneously capture intensional meaning and relational context. Empirical evaluation on a pizza restaurant relational dataset demonstrates that the approach produces diverse, semantically coherent concept names, effectively supporting concept interpretation, knowledge validation, and diagnostic assessment of symbolic data quality.
This study investigates the formation mechanism of shared conceptual spaces in multilingual large language models and their impact on cross-lingual transfer, which remain poorly understood. Focusing on the EuroLLM pretraining dynamics, the authors combine activation patching with cross-lingual concept representation isolation to trace the emergence and evolution of language-agnostic representations. They further assess the causal influence of these representations on translation behavior by injecting translation prompts. The findings reveal that a shared conceptual space forms early in training and undergoes continuous refinement, with its alignment quality exhibiting language-specific variation. Notably, observed improvements in translation performance often stem from shifts in decoding strategies—such as altered word-sense selection or handling of homographs—rather than genuine gains in cross-lingual capability. This work offers novel insights into the training dynamics and causal interpretability of multilingual models.
Direct application of large language models (LLMs) to taxonomy expansion often introduces noise, redundancy, and hierarchical inconsistencies, undermining the reliability of automated extension. To address this, this work proposes ReLTEx, a novel framework that synergistically integrates LLM-generated candidate concepts with a structure-aware validation mechanism and a recursive expansion control strategy. This design effectively mitigates model hallucinations while preserving semantic coherence and hierarchical consistency in the expanded taxonomies. Through a comprehensive evaluation protocol combining masked assessment, tailored metrics, and human validation, experiments demonstrate that ReLTEx significantly outperforms existing methods across multiple benchmark taxonomies, achieving substantial improvements in both reliability and semantic fidelity.
This study investigates the reliability and limitations of large language models (LLMs) in automatically generating entity-relationship (ER) diagrams from complex natural language requirements. Employing prompt strategies including zero-shot, chain-of-thought (CoT), and CoT augmented with a verifier, the authors systematically evaluate three leading LLMs on their ability to extract entities, relationships, and attributes from textual descriptions and produce conceptually consistent ER diagrams. The results indicate that while models perform adequately on low-complexity specifications, their outputs frequently suffer from logical inconsistencies, semantic ambiguities, and failures to correctly express constraints as requirement complexity increases. The findings highlight fundamental shortcomings of current LLMs in high-stakes database modeling tasks and provide empirical evidence for the role of prompt engineering in structured conceptual modeling.