Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification

📅 2026-07-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing graph construction methods based on feature similarity often introduce semantically irrelevant adjacency relationships, which hinder the performance of semi-supervised image classification. This work proposes a novel approach that leverages large language models (LLMs) to refine the semantic structure of image graphs. Specifically, it first employs a vision-language model together with an LLM to generate textual descriptions for images, then utilizes the LLM to assess pairwise semantic similarity between images. These similarity scores are used to reweight edges in both k-nearest neighbor (kNN) and mutual kNN graphs, followed by text-guided edge pruning to enhance semantic consistency within the graph. The resulting semantically refined graphs significantly improve classification accuracy of graph convolutional networks across multiple backbone architectures, with particularly pronounced gains observed in kNN-based graphs.
📝 Abstract
While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and unlabeled data, have emerged as a promising solution. One of the primary challenges in applying GCNs to image classification is graph construction, since, unlike in citation networks or similar domains, images typically do not come with a predefined structural representation. For visual data, most studies construct graphs based on the similarity between feature vectors from pretrained deep learning backbones, typically by employing kNN or reciprocal kNN algorithms. Although Large Language Models (LLMs) have shown remarkable capability in capturing high-level semantics, their integration with GCNs for image classification remains underexplored. Aiming to fill this gap, our approach uses a Vision Language Model (VLM) to generate textual image descriptions, which are then processed by an LLM to estimate semantic similarity scores between connected images. These scores guide the pruning of edges in kNN and reciprocal kNN graphs, filtering out semantically irrelevant neighbors. Experimental results reveal that leveraging LLMs for graph refinement can improve classification accuracy, particularly for kNN graphs and some backbones. The source code is publicly available at http://gcnllm.lucasvalem.com.
Problem

Research questions and friction points this paper is trying to address.

Semi-Supervised Image Classification
Graph Convolutional Networks
Large Language Models
Graph Construction
Semantic Similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Graph Convolutional Networks
Semi-Supervised Learning
Vision-Language Models
Graph Construction
🔎 Similar Papers
No similar papers found.
C
Camila Piscioneri Magalhães
Institute of Mathematics and Computer Science (ICMC), University of São Paulo (USP), São Carlos – SP – Brazil
L
Lucas Pascotti Valem
Institute of Mathematics and Computer Science (ICMC), University of São Paulo (USP), São Carlos – SP – Brazil