🤖 AI Summary
Graph neural networks (GNNs) for image data are typically constrained by predefined topologies—e.g., regular grids or hand-crafted superpixels—that fail to capture semantically meaningful pixel relationships. To address this, we propose a data-driven dynamic graph construction method based on pixel intensity correlations: specifically, intra-row, intra-column, and row–column cross-product correlations. This approach yields adaptive, semantically informed graph structures that jointly model local and global pixel dependencies without relying on manually designed topologies. We evaluate our method on MNIST and Fashion-MNIST using representative GNN architectures—including Graph CNN, GAT, and GatedGCN—and demonstrate consistent improvements in classification accuracy. Our learned graph structures outperform conventional grid-based and superpixel-based baselines by an average of 1.2–2.8 percentage points, confirming that data-adaptive graph construction significantly enhances the representational capacity of image GNNs.
📝 Abstract
Image datasets such as MNIST are a key benchmark for testing Graph Neural Network (GNN) architectures. The images are traditionally represented as a grid graph with each node representing a pixel and edges connecting neighboring pixels (vertically and horizontally). The graph signal is the values (intensities) of each pixel in the image. The graphs are commonly used as input to graph neural networks (e.g., Graph Convolutional Neural Networks (Graph CNNs) [1, 2], Graph Attention Networks (GAT) [3], GatedGCN [4]) to classify the images. In this work, we improve the accuracy of downstream graph neural network tasks by finding alternative graphs to the grid graph and superpixel methods to represent the dataset images, following the approach in [5, 6]. We find row correlation, column correlation, and product graphs for each image in MNIST and Fashion-MNIST using correlations between the pixel values building on the method in [5, 6]. Experiments show that using these different graph representations and features as input into downstream GNN models improves the accuracy over using the traditional grid graph and superpixel methods in the literature.