Score
Designs, constructs, and analyzes graph representations of image superpixels by segmenting images into coherent regions, extracting per-region features (e.g., frozen CNN descriptors), and creating nodes connected by edges that encode spatial adjacency and inter-regional geometry. Builds and adapts graph-transformer or graph neural network architectures that operate on this region-level graph for downstream tasks.
This work addresses the limitations of existing image classification approaches: convolutional neural networks (CNNs) and Vision Transformers (ViTs) incur high computational costs, while superpixel-based graph methods struggle to model long-range dependencies and lack geometric transformation invariance. To overcome these issues, the authors propose a novel superpixel graph classification framework that incorporates a geometrically invariant preprocessing step to preserve translation and rotation invariance, and introduces a Graph Transformer to effectively capture long-range dependencies. Notably, this is the first method to jointly achieve geometric invariance and long-range modeling on superpixel graphs without relying on superpixel boundary coordinates. Experimental results on CIFAR-10 demonstrate that the proposed approach significantly outperforms multiple baselines and attains accuracy comparable to the current state-of-the-art model, ShapeGNN.
This work addresses the information loss inherent in pixel aggregation within existing image classification methods by proposing the Superpixel Transformer (SPT) framework. SPT uniquely integrates superpixel representations, graph attention networks, and Vision Transformers into a unified architecture, supporting arbitrary superpixel segmentation strategies, graph connectivity patterns, and multidimensional sinusoidal positional encodings. By leveraging an enhanced superpixel data structure combined with self-attention mechanisms, the model preserves local structural details while significantly boosting representational capacity. Experimental results demonstrate that SPT substantially outperforms conventional superpixel-based graph neural network approaches on benchmarks including CIFAR-10, Fashion-MNIST, and Imagenette, achieving performance comparable to standard Vision Transformers. Furthermore, the study validates that constrained graph connectivity effectively enhances the modeling capability of Transformers.
Standard convolutional operations cannot be directly applied to graph-structured data due to its irregular, non-Euclidean topology. Method: This paper proposes Graph Kernel-driven Learnable Structural Convolution (GK-Conv), a purely structural, end-to-end modeling framework operating directly on non-Euclidean graph domains. GK-Conv eliminates explicit graph embedding and instead constructs a parameterized, structural convolutional operator grounded in generic graph kernel functions—enabling plug-and-play integration of arbitrary graph kernels and generating CNN-style, interpretable structural masks. The model is fully differentiable and optimized via ablation-guided hyperparameter analysis. Contribution/Results: GK-Conv achieves state-of-the-art performance across multiple graph classification and regression benchmarks, empirically validating the central claim that strong generalization can be attained using topology alone—without node or edge features.
Standard grid-based patching in Vision Transformers (ViTs) often yields tokens containing mixed semantic content, undermining representation consistency. To address this, we propose a superpixel-driven, semantically consistent tokenization method—the first to integrate superpixels into ViT token generation. Our approach features a two-stage pipeline: pre-aggregation feature extraction followed by superpixel-aware aggregation. It leverages SLIC superpixel segmentation, region-wise feature pre-aggregation, deformable attention adaptation, and token-level semantic alignment to resolve the compatibility challenge between irregular superpixel regions and the Transformer architecture. The method is backbone-agnostic and plug-and-play, requiring no modifications to the base network. Extensive experiments on ImageNet, CIFAR-100, and adversarial robustness benchmarks demonstrate significant improvements in both accuracy and generalization, empirically validating the critical performance gain conferred by semantically pure tokens in ViTs.
This paper introduces the first unconditional joint generation task of scene graphs and corresponding images, aiming to simultaneously synthesize structured scene graphs—comprising object categories, bounding boxes, and relational triplets—and photorealistic images from noise, enabling controllable and interpretable visual content generation. To this end, we propose DiffuseSG: a graph Transformer-based diffusion denoiser that unifies modeling of nodes (categories + coordinates), edges (relations), and adjacency matrices. We introduce IoU regularization and a continuous–discrete co-optimization mechanism, and pioneer the embedding of discrete category labels into a continuous latent space for joint diffusion modeling. Evaluated on Visual Genome and COCO-Stuff, DiffuseSG significantly outperforms state-of-the-art methods in both joint generation quality and fidelity. Moreover, it improves downstream scene graph completion and object detection performance, and generates high-fidelity samples that enhance model training through data augmentation.
Existing deep learning models struggle to effectively encode spatial, topological, and semantic structural information inherent in images. This work systematically evaluates the impact of various visual graph construction strategies on image classification performance within a unified three-layer Graph Convolutional Network (GCN) framework. For the first time, it demonstrates that the graph structure itself plays a decisive role in model performance. The study underscores the critical importance of the graph construction preprocessing stage, providing empirical evidence that well-designed graph structures substantially enhance classification accuracy. These findings offer both methodological guidance and practical justification for graph structure selection and preprocessing in visual graph neural networks.
This work addresses the challenges posed by high structural heterogeneity, large intra-class variation, and subtle visual differences between benign and malignant lesions in dermoscopic images. To this end, the authors propose a superpixel-based multimodal fusion approach that models lesions as graphs whose nodes correspond to superpixels. Node features are extracted using a frozen CNN, while geometric relationships are incorporated as edge attributes. A novel metadata context node is introduced to enable native graph-level fusion of clinical information with visual features. Discriminative classification embeddings are generated through an edge-aware Graph Transformer coupled with an attention propagation mechanism. This method, which uniquely integrates superpixel graph structure, geometric edge attributes, and metadata context, achieves state-of-the-art performance across four public datasets, significantly improving both accuracy and robustness in benign–malignant skin lesion classification.
This work addresses the limitations of existing superpixel methods, which produce irregular regions that are misaligned with regular operators such as convolutions, thereby hindering parallel computation and end-to-end deep learning. To overcome this, the study introduces granular ball computing into superpixel generation for the first time, proposing a structured superpixel representation based on multi-scale square blocks. By evaluating pixel intensity similarity to compute purity scores, the method adaptively selects high-quality square blocks for image coverage. This formulation inherently supports efficient parallel processing and integrates seamlessly into graph neural networks (GNNs) or Vision Transformers (ViTs) for end-to-end training. Experiments across multiple downstream vision tasks demonstrate that the proposed square superpixels significantly enhance performance, validating their advantages in both structured representation and computational efficiency.
Existing visual graph recognition methods are often confined to specific tasks and lack generalizability and cross-scenario transferability. This work proposes GraSP, an end-to-end framework based on subgraph prediction that jointly models graph structure and visual features to enable unified recognition of diverse graph types and rendering styles. GraSP achieves cross-task transfer without task-specific fine-tuning, representing the first general-purpose and transferable approach for visual graph recognition. Evaluated on multiple synthetic benchmarks and a real-world application, GraSP demonstrates exceptional generalization and adaptability, advancing the field toward a unified paradigm for graph recognition.
This work addresses the challenge of semi-supervised image classification under limited labeled data by proposing a novel approach that integrates multi-source features with graph-structured representations. The method constructs a multi-graph representation by fusing complementary features extracted from convolutional neural networks (CNNs), Vision Transformers (ViTs), and graph convolutional networks (GCNs), and refines the graph structure through manifold learning. A key innovation is the introduction of a ranking-based aggregation mechanism to effectively integrate heterogeneous feature sources. Extensive experiments demonstrate that the proposed framework consistently achieves significant improvements in classification accuracy across various settings, thereby validating the efficacy of the multi-feature fusion strategy and the graph optimization scheme.