Score
Design and build graph-structured representations from images by segmenting an image into nodes (pixels, superpixels, or regions) and defining edges using spatial, topological, or similarity/neighborhood relations; encode node and edge attributes that capture visual and semantic features and output adjacency/feature formats compatible with graph algorithms and graph neural networks. Techniques include superpixel or region generation, k-NN/epsilon neighborhood graphs, spatial/topological adjacency, and learned or attribute‑based edge construction.
Existing deep learning models struggle to effectively encode spatial, topological, and semantic structural information inherent in images. This work systematically evaluates the impact of various visual graph construction strategies on image classification performance within a unified three-layer Graph Convolutional Network (GCN) framework. For the first time, it demonstrates that the graph structure itself plays a decisive role in model performance. The study underscores the critical importance of the graph construction preprocessing stage, providing empirical evidence that well-designed graph structures substantially enhance classification accuracy. These findings offer both methodological guidance and practical justification for graph structure selection and preprocessing in visual graph neural networks.
This paper introduces the first unconditional joint generation task of scene graphs and corresponding images, aiming to simultaneously synthesize structured scene graphs—comprising object categories, bounding boxes, and relational triplets—and photorealistic images from noise, enabling controllable and interpretable visual content generation. To this end, we propose DiffuseSG: a graph Transformer-based diffusion denoiser that unifies modeling of nodes (categories + coordinates), edges (relations), and adjacency matrices. We introduce IoU regularization and a continuous–discrete co-optimization mechanism, and pioneer the embedding of discrete category labels into a continuous latent space for joint diffusion modeling. Evaluated on Visual Genome and COCO-Stuff, DiffuseSG significantly outperforms state-of-the-art methods in both joint generation quality and fidelity. Moreover, it improves downstream scene graph completion and object detection performance, and generates high-fidelity samples that enhance model training through data augmentation.
Existing graph generation methods neglect edge attribute modeling, limiting their applicability in domains such as transportation that require rich edge features. To address this, we propose the first score-based diffusion framework jointly modeling nodes, edges, and adjacency structure. Our approach introduces a novel node-edge joint attention mechanism that enables bidirectional dependency modeling among all three components throughout the diffusion process, supporting high-fidelity edge attribute generation. Key technical innovations include score distillation, edge-aware diffusion sampling, and joint noise modeling over node and edge variables. Extensive experiments on multiple real-world and synthetic benchmarks—including a newly constructed edge-valued evaluation dataset—demonstrate significant improvements over state-of-the-art methods: 21.3% reduction in edge attribute reconstruction error and 18.7% gain in structural validity. The method has been successfully deployed for traffic scenario graph generation.
Existing image segmentation methods predominantly rely on pixel-wise losses (e.g., Dice), neglecting topological consistency; mainstream topology-aware approaches either lack rigorous theoretical guarantees or suffer from high computational cost and poor generalizability. Method: We propose the first differentiable, lightweight, and formally guaranteed topology-preserving framework: (i) we introduce and optimize a strict homotopy equivalence metric; (ii) we construct a differentiable component graph based on connected components, enabling local neighborhood-sensitive topological modeling and loss computation; (iii) we employ graph neural networks for feature aggregation and homotopy classification to enforce homotopy equivalence between predictions and ground truth. Contribution/Results: Our method achieves state-of-the-art performance on diverse multi-class medical and natural image segmentation benchmarks. It significantly improves topological accuracy and accelerates topological loss computation by 5× compared to persistent homology–based methods.
Existing image captioning datasets rely solely on unstructured text, failing to explicitly encode compositional structures and relational semantics among entities. To address this, we propose Graph-based Captioning (GBC), a novel paradigm where nodes represent entities, attributes, and relation phrases, while labeled edges explicitly model semantic connections—preserving linguistic flexibility while introducing hierarchical structure. We introduce the first graph-based annotation framework, enabling automatic construction of the large-scale GBC10M dataset (10 million samples). Moreover, we pioneer the use of graph structure as both a supervision signal in CLIP-style contrastive learning and an intermediate representation for text-to-image generation. Integrating multimodal large models, object detection, and graph modeling, our approach achieves significant improvements across VQA, referring expression comprehension (REC), and captioning benchmarks. Experiments demonstrate that graph-structured representations enhance both fidelity and fine-grained controllability in text-to-image synthesis. Code and the GBC10M dataset are publicly released.
This work addresses the challenge of efficiently compressing edge weights in weighted graph adjacency matrices by proposing a line-graph-based graph signal modeling approach. Specifically, edge weights are treated as graph signals defined on the line graph and are compressed through transform coding using graph filter banks, followed by quantization and entropy coding. The method innovatively introduces an edge smoothness metric that can be computed without explicitly constructing the line graph, enabling effective prediction of compression performance. Experimental results demonstrate that the proposed framework consistently outperforms existing matrix preprocessing techniques on both synthetic and real-world datasets, thereby validating its efficacy and practicality for lossy graph weight compression.