KGATE : a Knowledge Graph Embedding Training Environment

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing knowledge graph embedding libraries, including the lack of autoencoder support, insufficient maintenance, and incomparable results. To overcome these issues, this work proposes a modular Python framework built upon PyTorch Geometric and TorchKGE. The proposed approach enables flexible end-to-end assembly of encoders and decoders, integrates essential components such as initializers, loss functions, and negative sampling strategies, and incorporates built-in preprocessing mechanisms to prevent data leakage. Benchmark evaluations demonstrate that the framework achieves training efficiency comparable to state-of-the-art baselines while offering more comprehensive functionalities. Consequently, it significantly enhances cross-library experimental reproducibility and the reliability of model comparisons.
📝 Abstract
Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction. Most KGE models follow an autoencoder architecture, in which an encoder projects the knowledge graph into the latent space and a decoder reconstruct it. Combining both encoder and decoder components is increasingly needed, yet existing libraries rarely support complete autoencoders, are often unmaintained, rely on undocumented default hyperparameters, and produce results that cannot be compared across libraries. Here we present KGATE (Knowledge Graph Autoencoder Training Environment), a modular Python library built on PyTorch Geometric and TorchKGE. KGATE lets users assemble initializers, encoders, decoders, losses, negative samplers, and evaluation metrics as building blocks, or plug in their own block. KGATE includes a preprocessing procedure that controls data leakage, a builtin training pipeline, and reproducibility by design. Benchmarks against six existing KGE libraries show that KGATE training time is comparable with the fastest libraries while offering a broader set of features.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Graph Embedding
Autoencoder
Reproducibility
Benchmarking
Library
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Graph Embedding
Autoencoder
Modular Library
Reproducibility
PyTorch Geometric
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Benjamin Loire
Aix Marseille Univ., INSERM, Marseille Medical Genetics, Systems Biomedicine Team, Marseille, France
G
Galadriel Brière
Aix Marseille Univ., INSERM, Marseille Medical Genetics, Systems Biomedicine Team, Marseille, France
C
Célia Brahimi
Aix Marseille Univ., INSERM, Marseille Medical Genetics, Systems Biomedicine Team, Marseille, France
A
Antoine Toffano
LIRMM, Univ. Montpellier, CNRS, Montpellier, France
Anaïs Baudot
Anaïs Baudot
CNRS - INSERM - Aix*Marseille Université
Systems and Networks BiologyBioinformaticsComputational Biology