Score
Designs and implements methods to represent and integrate semantic relationships among labels or concepts into models and inputs; builds encoders, prompt mechanisms, or architectural components that capture inter-class topology and hierarchical correlations so label dependencies are preserved and exploitable.
Real-world knowledge often involves high-order n-ary relations, yet conventional knowledge graphs (KGs) flatten them into binary triples—causing semantic loss—while hypergraph models ignore role distinctions among entities within hyperedges. Method: We propose the first orthogonal two-dimensional taxonomy: the horizontal axis categorizes methodologies (e.g., translation-based, tensor decomposition, neural networks), and the vertical axis hierarchically classifies role-awareness granularity (role-agnostic, position-aware, role-aware), unifying modeling paradigms for knowledge hypergraphs and hyper-relational KGs. Our framework integrates KG principles, hypergraph theory, tensor decomposition, deep neural networks, logical rules, and hyperedge expansion techniques, covering the full pipeline—from data construction and negative sampling to evaluation. Contribution/Results: We systematically survey methodological evolution, benchmark datasets, and evaluation protocols; identify critical challenges—including insufficient role modeling and weak interpretability—and provide a structured research guide and paradigmatic framework for n-ary knowledge representation learning.
Addressing the challenges of automated domain ontology mapping, high human dependency, and excessive costs in semantic modeling of structured data (CSV/JSON/XML), this paper proposes the Knowledge Prompt Chaining (KPC) framework. KPC serializes graph-structured domain knowledge and injects it into large language models (LLMs) to enable structure-aware, end-to-end semantic annotation and knowledge graph generation. By integrating a prompt chaining architecture, graph-knowledge serialization, and lightweight LLM fine-tuning, KPC substantially reduces reliance on large-scale input data. Experimental results demonstrate that KPC outperforms state-of-the-art methods in both semantic annotation accuracy and knowledge graph quality. It establishes an efficient, scalable paradigm for semantic enrichment of structured data—particularly suitable for low-resource settings—while preserving fidelity to domain semantics and structural constraints.
Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.
Existing approaches lack formal mechanisms for semantically preserving structural transformations across heterogeneous representation systems (e.g., formal languages, geometric diagrams, informal notations), especially for arbitrary user-specified semantic relations such as equivalence. Method: This paper introduces a representation-system-agnostic (RS-agnostic) structural migration calculus framework grounded in formal inference rules and pattern-based encoding. It enables verifiable, semantics-preserving structural mapping and transformation among disparate representation systems by integrating representation-system theory with constructive space modeling. Contributions: (1) The first formally verified RS-agnostic transformation calculus proven to satisfy arbitrary target semantic relations; (2) A pattern-driven, information-preserving mechanism supporting automatic, meaning-preserving reconstruction across multimodal representations; (3) A rigorous formal foundation for cross-representational cognitive modeling and intelligent representation generation. The framework achieves high generality in abstract representation transformation while ensuring semantic fidelity and verifiability.
This work addresses the limitation that conceptual knowledge in large language models (LLMs) is implicitly encoded through statistical correlations, lacking explicit, structured, and composable representations—hindering model stability, controllability, and alignment with human cognition. To tackle this, the paper introduces the first design-space taxonomy for conceptual structures in LLMs, proposing a “four-stage–two-source” framework that spans training, architecture, inference, and interpretation, combined with internal derivation and external anchoring. The framework systematically integrates probing, dictionary learning, and external knowledge-guided approaches, revealing critical gaps in current research—particularly insufficient exploration of the inference stage, fragmentation across stages, and inconsistent terminology. By shifting the paradigm from passive discovery to active design of conceptual representations, this study provides a theoretical foundation and strategic direction for developing LLMs that are more controllable, interpretable, and cognitively aligned.
Semantic communication faces compatibility challenges with operational networks designed around network-centric metrics, violating the separation-of-concerns principle. Method: This work proposes a task-driven semantic communication service architecture that jointly rethinks representation, interface, and control. It introduces (i) standardized non-arbitrary semantic encoding, (ii) an application-network co-designed standard interface, and (iii) a semantic-task-oriented dedicated control plane, along with a novel text semantic transmission protocol. Contribution/Results: Evaluated across three representative text transmission scenarios, the architecture significantly improves semantic fidelity and task completion efficiency. It constitutes the first deployable paradigm enabling native semantic communication support in operational networks—bridging the gap between semantic communication theory and real-world network deployment.
This study addresses the challenge of agents adapting to unknown semantic constraints and assembling novel structures post-deployment. We propose a neuro-symbolic architecture integrating natural language with visual observations, utilizing dialogue and demonstrations as symbolic evidence to facilitate online constraint acquisition and dynamic structure assembly. Experiments in a simulated truck assembly task demonstrate that language guidance significantly enhances online adaptive learning efficiency compared to methods relying solely on demonstrations or component naming. By validating the efficacy of neuro-symbolic reasoning in open-world structured tasks, this work establishes a novel paradigm for continuous agent learning, enabling robust adaptation to previously unseen semantic requirements through multimodal symbolic grounding.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This work addresses the limitations of existing concept-based models, which rely on fine-grained annotations and treat concepts as flat, independent units, thereby hindering the construction of interpretable, hierarchical concept structures. To overcome this, the authors propose Multi-Level Concept Segmentation (MLCS) and Deep Hierarchical Concept Embedding Models (Deep-HiCEMs), which require only coarse-grained top-level supervision to automatically discover multi-layered, human-interpretable concept hierarchies. The framework supports concept interventions across abstraction levels and successfully uncovers novel, explainable concepts absent from training data across multiple benchmarks. While maintaining high predictive accuracy, the method significantly enhances task performance through test-time interventions, marking the first approach capable of automatically constructing a multi-granular concept system from coarse-grained labels alone.
This study addresses the limitations of traditional Design Structure Matrix (DSM) modularization approaches, which rely solely on graph-based optimization and lack engineering semantic context, often failing to align with practical design requirements. The authors propose a novel DSM modularization paradigm integrating large language models (LLMs), leveraging prompt engineering and iterative refinement to embed system-level semantic information directly into the partitioning process—achieving high-quality results without custom optimization code. Central to this work is the "semantic alignment hypothesis," which elucidates how improper incorporation of domain knowledge can degrade performance. Through systematic experiments across five representative engineering cases using three mainstream LLMs, the method demonstrates convergence to reference-quality modularization within 30 iterations, offering a reproducible and practical pathway for LLM-driven engineering design optimization.
研究通过提出表示赋能方法,解决在持续模型构建中决定哪些元素值得被表示的问题,以提高模型的未来建模和规划能力。