Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification

📅 2026-05-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the information loss inherent in pixel aggregation within existing image classification methods by proposing the Superpixel Transformer (SPT) framework. SPT uniquely integrates superpixel representations, graph attention networks, and Vision Transformers into a unified architecture, supporting arbitrary superpixel segmentation strategies, graph connectivity patterns, and multidimensional sinusoidal positional encodings. By leveraging an enhanced superpixel data structure combined with self-attention mechanisms, the model preserves local structural details while significantly boosting representational capacity. Experimental results demonstrate that SPT substantially outperforms conventional superpixel-based graph neural network approaches on benchmarks including CIFAR-10, Fashion-MNIST, and Imagenette, achieving performance comparable to standard Vision Transformers. Furthermore, the study validates that constrained graph connectivity effectively enhances the modeling capability of Transformers.
📝 Abstract
Superpixel-based image classification has traditionally leveraged graph neural networks (GNNs) for processing irregular image representations. Recent advances in computer vision, driven by Vision Transformers (ViTs), have introduced new paradigms in self-attentional models, surpassing convolutional neural networks (CNNs) in various tasks. However, a synergistic connection between GNNs, superpixels, and transformers remains unexplored. In this work, we propose Superpixel Transformers (SPT), a novel framework that unifies superpixel-based image classification and ViTs. SPT generalizes the Superpixel Image Classification with Graph Attention Networks (SICGAT) model and ViT to support arbitrary superpixel-based chunking strategies, connectivity graphs, and positional encodings. We introduce refinements including a multidimensional sine-cosine positional encoding and an enriched patch data structure that fully incorporates superpixel shape and color information. By testing SPT across datasets such as CIFAR10, FashionMNIST, and Imagenette, with various superpixel generation and graph connectivity strategies, we demonstrate that SPT achieves superior performance compared to previous superpixel-based GNN methods and remains competitive with ViTs. Notably, our approach addresses the limitations of SICGAT, such as information loss during pixel aggregation, and shows how constrained graph connectivity can enhance ViT performance. SPT bridges the gap between superpixel-based and transformer models, opening avenues for cross-domain generalization and future innovations in hybrid attentional frameworks, and showing that an image can also be worth $16\times16$ superpixels.
Problem

Research questions and friction points this paper is trying to address.

superpixel
Vision Transformers
graph neural networks
image classification
self-attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Superpixel Transformers
Vision Transformers
Graph Neural Networks
Positional Encoding
Image Classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pedro Henrique da Costa Avelar
Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), 91501-970, Porto Alegre, Rio Grande do Sul, Brazil; Division of Informatics, School of Health Sciences, Imaging and Data Science, Faculty of Biology, Medicine and Health, University of Manchester, M13 9GB, Vaughan House, Portsmouth St, Manchester, United Kingdom
A
Anderson R. Tavares
Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), 91501-970, Porto Alegre, Rio Grande do Sul, Brazil
L
Luís C. Lamb
Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), 91501-970, Porto Alegre, Rio Grande do Sul, Brazil