Gromov-Wasserstein Distillation for Inductive Multi-View Embedding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of Gromov-Wasserstein multidimensional scaling (GW-MDS), which is inherently transductive and cannot generalize to unseen samples. To overcome this, we propose an inductive learning framework based on barycentric distillation. This method establishes, for the first time, a bridge between transductive GW embeddings and inductive neural networks via barycentric projection. Specifically, a teacher model generates target-space representations, and knowledge distillation is employed to train a student network that learns an explicit out-of-sample mapping while supporting multi-view consensus learning. Experiments on both synthetic and real-world datasets demonstrate that the proposed framework effectively preserves data geometric structures, achieving significantly superior performance compared to direct neural GW training baselines.
📝 Abstract
Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing no explicit mapping for unseen samples. We introduce an inductive framework based on barycentric distillation. A GW-MDS teacher learns a latent support and an optimal transport plan from the training data, and barycentric projection converts the resulting coupling into sample-aligned targets. A neural student then learns an explicit out-of-sample mapping, avoiding additional relational-matrix construction and GW optimization at inference. We formulate the approach for single-view data and extend it to Mean-GWMDS and Multi-GWMDS teachers through consensus and selected-projection targets learned by a multi-view student with view-specific encoders. We also investigate a direct neural baseline trained solely with a GW objective. Experiments on synthetic and real-world data using Euclidean, geodesic, and cosine relations show that the distilled models preserve the teacher geometry on unseen samples and consistently outperform direct neural GW training in sample-indexed relational preservation. These results establish barycentric projection as an effective bridge between transductive GW embeddings and inductive neural mappings.
Problem

Research questions and friction points this paper is trying to address.

Gromov-Wasserstein
inductive embedding
multi-view learning
out-of-sample mapping
distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gromov-Wasserstein Distillation
Barycentric Projection
Inductive Multi-View Embedding
Optimal Transport
Teacher-Student Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.