Score
Design and build structured graph representations that model entities as nodes and their relations and attributes as edges and node features, including entity-centric and operational-entity graphs. Use these representations to transform unstructured observations into standardized features, encode relational connectivity and competition-like relations, and support analysis or learning with relational graph convolutional and attention-based algorithms.
End-to-end representation learning for multi-table relational databases remains challenging due to the need for manual feature engineering and the lack of unified structural abstractions across heterogeneous, time-evolving tables. Method: We propose Relational Deep Learning (RDL), a framework that bypasses traditional feature engineering by formally introducing the *temporal heterogeneous relational entity graph*—a unified graph structure where primary–foreign key relationships define edges, schema constraints govern node/edge types, and timestamps serve as dynamic attributes. RDL integrates graph neural networks (GNNs), relational algebraic modeling, temporal graph learning, and heterogeneous graph architectures to enable cross-table joint representation learning. Contribution: This work establishes the first theoretical foundation and technical roadmap for RDL, systematically identifying core challenges and curating benchmark datasets. It advances graph representation learning toward relational data foundation models and introduces a novel paradigm for large-scale, multi-table joint modeling.
This study addresses the poorly understood synergy between graph embedding and entity linking in entity retrieval. We systematically evaluate the combinatorial effects of three graph embedding models and five entity linking methods, and propose an embedding-distance-based re-ranking strategy for retrieval. To our knowledge, this is the first work to quantitatively analyze their joint impact: we find that joint embeddings—integrating both graph structure and textual descriptions—significantly outperform unimodal embeddings; moreover, entity linkers must jointly optimize concept-level precision and recall, rather than maximizing either in isolation. Experiments show that final retrieval performance depends critically on both graph embedding choice and linker capability—especially on high-coverage knowledge graphs. Our core contribution lies in revealing the pivotal role of cross-modal modeling and balanced linking for retrieval quality, providing a principled, interpretable, and reproducible methodology for entity retrieval. (149 words)
Large language models (LLMs) struggle to effectively comprehend graph-structured data due to their inherent sequence-based architecture and lack of native graph-aware representations. Method: This paper introduces *graph laws*—statistically derived, topologically parameterized features that are interpretable as natural language descriptions—establishing a novel paradigm for representing graphs as LLM-compatible inputs. We systematically construct a multi-dimensional graph law framework spanning macro/micro scales, low/high orders, and static/dynamic properties, integrating graph-theoretic analysis, multi-scale observational modeling, and natural language alignment techniques, while establishing semantic mappings to downstream graph tasks and retrieval-augmented generation (RAG) scenarios. Results: Experiments demonstrate that graph laws substantially mitigate LLM hallucination, overcome context-length limitations, and enable end-to-end graph reasoning. The approach achieves strong generalization across diverse domains, including molecular design, recommender systems, and protein structure modeling.
Existing knowledge graph reasoning methods suffer from limited performance in entity classification and link prediction due to static neighbor aggregation and insufficient semantic modeling. To address this, we propose a dynamic representation learning framework that integrates attention mechanisms with graph convolutional networks (GCNs). Our key contributions are: (1) the first incorporation of attention mechanisms into every GCN layer to enable relation-aware, adaptive neighbor weighting during aggregation; and (2) a guided node representation learning paradigm based on entity similarity, which jointly models attribute and relational information to enhance implicit semantic capture. Extensive experiments on standard knowledge graph benchmarks demonstrate that our method achieves an average accuracy improvement of over 5.2% compared to state-of-the-art GNNs and embedding models on both entity classification and link prediction tasks. The results confirm significant gains in fine-grained entity representation quality and reasoning generalizability.
This work addresses the limitation of existing graph learning approaches, which typically operate in isolation within a single modality and task, thereby hindering the cross-task and cross-modal reuse of structural knowledge. To overcome this, the authors propose G-Substrate, a novel framework that models graph structures as persistent, shareable substrates. By unifying structural patterns and employing a role-interleaved training strategy, G-Substrate enables collaborative learning across multiple tasks and modalities. This approach facilitates the continuous accumulation and transfer of graph-structured knowledge, consistently outperforming both isolated training and conventional multi-task learning methods across diverse domains, modalities, and tasks.
This work addresses the challenge that graph structures naively derived from relational databases often suffer from information overload and semantic fragmentation, rendering them ill-suited for relational reasoning with graph neural networks (GNNs). To overcome this limitation, the authors propose an end-to-end structure optimizer that automatically constructs GNN-friendly relational graphs by jointly performing information filtering and semantic enrichment. The method uncovers key mechanisms underlying effective adaptation of relational graphs to GNN architectures. Empirical evaluation across 26 diverse tasks—including classification, regression, and recommendation—demonstrates consistent improvements in model accuracy, frequently accompanied by reduced inference overhead.
Recent progress in language modeling has expanded the range of tasks that can be approached through natural language interfaces, including problems that require structured reasoning. However, it remains unclear how effectively limited-capacity language models can infer formal properties of relational structures when those structures are presented in textual form. Understanding the conditions under which structured reasoning succeeds or fails is essential for applying small models in graph-based domains. We conduct a systematic study of graph-theoretic property inference in small instruction-tuned language models, isolating the roles of input representation and reasoning strategy. Across a diverse set of local and global graph metrics, we find that structural performance is highly sensitive to how relational information is organized. Representations that preserve neighborhood structure consistently improve estimation stability and ordinal consistency, while multi-branch reasoning yields the most reliable aggregate gains across configurations. These results show that graph property inference in small language models depends critically on representational organization and inference design. Structural competence is therefore shaped not only by model scale, but by how relational information is encoded and how predictions are elicited. The findings identify practical levers for improving structured inference under constrained model capacity.
This work addresses the limitations of traditional deep learning approaches for relational databases, which rely on manual feature engineering and often fail to preserve critical relational structures, as well as the restricted generalization capability of existing relational deep learning models. To overcome these challenges, the authors propose a lightweight hybrid architecture that effectively integrates a fine-tuned BART encoder with a GraphSAGE graph neural network for the first time: the BART component captures intra-row semantic information, while GraphSAGE propagates representations over an entity-relation graph to inject structural context. By jointly modeling semantics and relational dependencies, the method achieves a ROC-AUC of 67.40 on the driver-dnf task in RelBench—approaching the performance of LightGBM (68.86) and specialized relational deep learning approaches (72.62)—demonstrating its effectiveness and strong generalization potential.
This work proposes FROG, a novel framework that addresses the limitations of existing relational deep learning approaches, which typically rely on fixed graph structures and struggle to adaptively optimize for improved predictive performance. FROG introduces table roles—derived from relational database schemas—into graph structure learning, constructing a learnable, full-resolution graph where tables can function as either nodes or edges in message passing. By integrating a role-driven message-passing mechanism with functional dependency constraints, FROG jointly optimizes both graph topology and GNN representations while preserving semantic consistency. Extensive experiments demonstrate that FROG significantly outperforms state-of-the-art models across multiple relational prediction tasks, highlighting the critical influence of table roles on downstream performance and establishing a new paradigm for graph construction in relational deep learning.
This work addresses the challenge of fine-grained similarity assessment in e-commerce entity search, where relevance depends on both product categories and contextual cues—a task poorly handled by conventional embedding methods due to their inability to model attribute correlations. The authors propose a two-stage zero-shot ranking approach: offline, a large language model (LLM) constructs a category-aware, structured attribute graph; online, a graph-augmented LLM efficiently ranks candidate entities by reasoning over this reusable graph. This is the first method to integrate a precomputed attribute graph with an LLM for zero-shot ranking without any training data. Experiments demonstrate that the approach outperforms baselines by over 5% in mean average precision, reduces per-item inference tokens by 57%, and exhibits strong cross-category generalization and practical deployment viability.