ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of entity semantic preservation, neighborhood interaction construction, and geometric prediction adaptation in multimodal graph foundation models by proposing a unified framework based on Clifford latent fields that integrates topological, textual, and visual information. Methodologically, it introduces node-indexed Cl(3) addressing and edge-aware geometric products to ensure cross-level semantic preservation, while combining Clifford algebraic operations with multi-depth structural propagation to construct an explicit latent field encoding architecture. Experimental results demonstrate that the proposed model achieves state-of-the-art performance across 11 datasets and 30 supervised and few-shot tasks, significantly outperforming existing baseline methods.
📝 Abstract
Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.
Problem

Research questions and friction points this paper is trying to address.

multimodal attributed graphs
graph foundation models
representation learning
entity semantics
interaction state
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clifford Latent Fields
Multimodal Graph Foundation Models
Geometric Algebra
Edge-aware Geometric Products
Task-Aligned Representation
🔎 Similar Papers
No similar papers found.