Score
Designs and implements encoders that convert scene graphs—nodes representing objects or regions and typed spatial or hierarchical edges—into fixed-size vector representations for downstream models. This includes computing spatial edges from maps, encoding hierarchical or branch structures and per-entity spatial context without inflating global token counts, and capturing geometric relationships and distributions among entities.
This work addresses the challenge of adapting vector geospatial data (points, lines, polygons) to conventional machine learning models. We propose a proximity encoding method based on multi-reference-point scaled distances, which losslessly maps geometric objects into continuous, centered, and type-uniform feature vectors—enabling, for the first time, unified vector encoding across geometric types. Our core innovations include spatial scaling normalization and parameterized vector embedding, jointly preserving shape centrality, geometric continuity, and high-fidelity spatial relationship modeling. Experiments demonstrate that the proposed encoding consistently outperforms rasterization-based baselines on both geometric discrimination and spatial relation identification tasks, yielding significant improvements in end-to-end vector geospatial AI modeling performance.
Existing node embedding methods suffer from two key limitations: vector addition lacks network semantic interpretability, and relationships among multi-scale (coarse-grained) embeddings remain ill-defined. This paper proposes a multi-scale node embedding framework that unifies the resolution of both issues for the first time. Leveraging a hierarchical coarse-graining mechanism grounded in renormalization theory, and imposing vector-sum constraints alongside low-dimensional reconstruction optimization in the embedding space, our method ensures that the embedding of any coarse-grained block node is strictly equal to the statistical mean of its constituent node embeddings. This guarantees statistical consistency across resolutions. Evaluated on international trade and input-output networks, the framework achieves high-fidelity structural reconstruction—e.g., accurate triangle counting—and supports arbitrary-scale graph generation. It significantly enhances interpretability and practicality in multi-scale graph modeling and synthesis.
This work addresses the challenge of efficiently compressing edge weights in weighted graph adjacency matrices by proposing a line-graph-based graph signal modeling approach. Specifically, edge weights are treated as graph signals defined on the line graph and are compressed through transform coding using graph filter banks, followed by quantization and entropy coding. The method innovatively introduces an edge smoothness metric that can be computed without explicitly constructing the line graph, enabling effective prediction of compression performance. Experimental results demonstrate that the proposed framework consistently outperforms existing matrix preprocessing techniques on both synthetic and real-world datasets, thereby validating its efficacy and practicality for lossy graph weight compression.
This work addresses the absence of rotation-equivariant positional encodings—specifically Rotary Position Embeddings (RoPE)—for non-grid graph-structured data. We propose WIRE, the first general-purpose RoPE extension to arbitrary graphs. WIRE constructs rotational transformations grounded in graph wavelet analysis, inherently satisfying permutation equivariance over node orderings and compatibility with linear attention mechanisms. Under mild conditions, it asymptotically approximates graph resistance distance, thereby explicitly encoding structural similarity. Parameter-free and plug-and-play, WIRE integrates seamlessly into diverse graph neural networks without architectural modification. Extensive experiments demonstrate that WIRE consistently outperforms existing positional encoding methods on tasks including monochromatic subgraph detection, point cloud semantic segmentation, and multiple standard graph benchmarks. Gains are especially pronounced on structure-sensitive tasks, validating both its theoretical foundation and empirical efficacy.
This paper introduces the first unconditional joint generation task of scene graphs and corresponding images, aiming to simultaneously synthesize structured scene graphs—comprising object categories, bounding boxes, and relational triplets—and photorealistic images from noise, enabling controllable and interpretable visual content generation. To this end, we propose DiffuseSG: a graph Transformer-based diffusion denoiser that unifies modeling of nodes (categories + coordinates), edges (relations), and adjacency matrices. We introduce IoU regularization and a continuous–discrete co-optimization mechanism, and pioneer the embedding of discrete category labels into a continuous latent space for joint diffusion modeling. Evaluated on Visual Genome and COCO-Stuff, DiffuseSG significantly outperforms state-of-the-art methods in both joint generation quality and fidelity. Moreover, it improves downstream scene graph completion and object detection performance, and generates high-fidelity samples that enhance model training through data augmentation.
This work addresses the challenge faced by visually impaired users in accessing node-link diagrams commonly distributed as bitmap images, a task for which existing assistive technologies are ill-suited due to their reliance on structured data rather than visual input. The paper presents the first lightweight deep learning approach for semantic segmentation of such diagram images, training a compact model on a large-scale synthetic dataset to achieve pixel-level parsing. The proposed method attains over 93% pixel accuracy on synthetic data and demonstrates strong performance both quantitatively and qualitatively. By enabling precise extraction of diagram semantics directly from rasterized images, this approach establishes a viable foundation for non-visual interaction and effectively bridges a critical gap in accessibility technology for bitmap-based graphical content.
This work addresses the challenge of efficient lossless compression for large-scale real-world graph data by proposing a novel algorithm that leverages geometric representations of graph structure through direct application of modern hyperbolic space embeddings. By capitalizing on the intrinsic hyperbolic geometry inherent in complex networks, the method achieves substantially improved compression efficiency while preserving lossless reconstruction. Experimental evaluation across diverse real-world graph datasets demonstrates that the proposed approach outperforms the current state-of-the-art methods by up to 42% in compression ratio, thereby validating the efficacy and superiority of hyperbolic embeddings for graph compression tasks.
This work addresses the challenge of modeling component assembly relationships in 2D scenes under few-shot settings without semantic labels. The authors propose an end-to-end scene graph generation method that integrates geometric feature extraction with structural reasoning. Specifically, Faster R-CNN is employed to extract geometric representations of components, and a Transformer architecture constructs an initial adjacency matrix. To refine relational inference, the approach incorporates an attention-based Graph Convolutional Network (aGCN) with a message-passing mechanism. Notably, the model operates without semantic supervision and achieves accurate recovery of ground-truth assembly relationships using only a minimal number of training samples. Experimental results on a toy vehicle dataset demonstrate the method’s effectiveness, significantly advancing the capability to model assembly relationships in semantically unlabeled scenarios.
Existing scene graph methods struggle to explicitly model the hierarchical entailment relations between locations and objects in Euclidean space, leading to insufficient structural consistency. This work proposes the first approach that incorporates hyperbolic geometry into scene graph representation learning, leveraging its innate capacity for encoding hierarchical structures. By integrating contrastive learning with attention mechanisms, the method enables more structured visual understanding. The resulting embeddings exhibit significantly improved hierarchical organization and semantic coherence, achieving a Graph IoU score of 33.51—representing an absolute gain of 8.14 over the strongest baseline—while maintaining competitive performance in retrieval tasks.