On Preserving Geometrical Invariance for Superpixel Image Classification using Graph Transformer

📅 2026-07-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing image classification approaches: convolutional neural networks (CNNs) and Vision Transformers (ViTs) incur high computational costs, while superpixel-based graph methods struggle to model long-range dependencies and lack geometric transformation invariance. To overcome these issues, the authors propose a novel superpixel graph classification framework that incorporates a geometrically invariant preprocessing step to preserve translation and rotation invariance, and introduces a Graph Transformer to effectively capture long-range dependencies. Notably, this is the first method to jointly achieve geometric invariance and long-range modeling on superpixel graphs without relying on superpixel boundary coordinates. Experimental results on CIFAR-10 demonstrate that the proposed approach significantly outperforms multiple baselines and attains accuracy comparable to the current state-of-the-art model, ShapeGNN.
📝 Abstract
Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information. Consequently, for a larger image dataset, CNNs and ViTs face deployability challenges due to high computational complexity. Representing images as graphs of superpixels offers an efficient alternative that preserves key information while eliminating pixel-level redundancy. Graph Neural Networks (GNNs) have been utilized on such graphs to perform image classification. However, GNNs are known to struggle with capturing long-range dependencies which is important in the domain of image classification. Furthermore, a majority of these superpixel-based image classification approaches do not explicitly preserve translation/rotation invariance. Nevertheless, preserving translation/rotation invariance is important for robust image classification. Thus, this paper proposes SuperGT, a Graph Transformer-based framework for image classification, which captures the long range dependencies, along with a pre-processing scheme that preserves translation/rotation invariance. We evaluate SuperGT on CIFAR-10 dataset and observe that it performs significantly better than many baselines. Furthermore, we note that the overall performance of SuperGT is comparable to the previous state-of-the-art model, namely, ShapeGNN, without relying on coordinates of the boundary points of each superpixel required by ShapeGNN.
Problem

Research questions and friction points this paper is trying to address.

superpixel
graph neural networks
long-range dependencies
geometrical invariance
image classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph Transformer
Superpixel
Geometrical Invariance
Image Classification
Long-range Dependencies
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sarabeshwar Balaji
Indian Institute of Science Education and Research Bhopal, Bhopal, Madhya Pradesh, 462066, India
S
Shubham Mohanty
Indian Institute of Science Education and Research Bhopal, Bhopal, Madhya Pradesh, 462066, India
A
Akash Anil
Indian Institute of Science Education and Research Bhopal, Bhopal, Madhya Pradesh, 462066, India