scene graph encoding

Designs and implements encoders that convert scene graphs—nodes representing objects or regions and typed spatial or hierarchical edges—into fixed-size vector representations for downstream models. This includes computing spatial edges from maps, encoding hierarchical or branch structures and per-entity spatial context without inflating global token counts, and capturing geometric relationships and distributions among entities.

scenegraphencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-Point Proximity Encoding For Vector-Mode Geospatial Machine Learning

Jun 05, 2025
JC
John Collins
🏛️ Odyssey Geospatial LLC

This work addresses the challenge of adapting vector geospatial data (points, lines, polygons) to conventional machine learning models. We propose a proximity encoding method based on multi-reference-point scaled distances, which losslessly maps geometric objects into continuous, centered, and type-uniform feature vectors—enabling, for the first time, unified vector encoding across geometric types. Our core innovations include spatial scaling normalization and parameterized vector embedding, jointly preserving shape centrality, geometric continuity, and high-fidelity spatial relationship modeling. Experiments demonstrate that the proposed encoding consistently outperforms rasterization-based baselines on both geometric discrimination and spatial relation identification tasks, yielding significant improvements in end-to-end vector geospatial AI modeling performance.

Captures geometric properties using reference point distancesEncodes vector-mode geospatial data for ML modelsImproves accuracy over rasterization-based methods

Multi-Scale Node Embeddings for Graph Modeling and Generation

Dec 05, 2024
RM
Riccardo Milocco
🏛️ IMT School for Advanced Studies | ING Bank N.V. | Leiden University

Existing node embedding methods suffer from two key limitations: vector addition lacks network semantic interpretability, and relationships among multi-scale (coarse-grained) embeddings remain ill-defined. This paper proposes a multi-scale node embedding framework that unifies the resolution of both issues for the first time. Leveraging a hierarchical coarse-graining mechanism grounded in renormalization theory, and imposing vector-sum constraints alongside low-dimensional reconstruction optimization in the embedding space, our method ensures that the embedding of any coarse-grained block node is strictly equal to the statistical mean of its constituent node embeddings. This guarantees statistical consistency across resolutions. Evaluated on international trade and input-output networks, the framework achieves high-fidelity structural reconstruction—e.g., accurate triangle counting—and supports arbitrary-scale graph generation. It significantly enhances interpretability and practicality in multi-scale graph modeling and synthesis.

Clarifying the network meaning of vector addition in embeddingsDeveloping consistent multiscale embeddings for network modeling and generationUnderstanding relationships between embeddings at different hierarchical scales

This work addresses the challenge of efficiently compressing edge weights in weighted graph adjacency matrices by proposing a line-graph-based graph signal modeling approach. Specifically, edge weights are treated as graph signals defined on the line graph and are compressed through transform coding using graph filter banks, followed by quantization and entropy coding. The method innovatively introduces an edge smoothness metric that can be computed without explicitly constructing the line graph, enabling effective prediction of compression performance. Experimental results demonstrate that the proposed framework consistently outperforms existing matrix preprocessing techniques on both synthetic and real-world datasets, thereby validating its efficacy and practicality for lossy graph weight compression.

edge weightsgraph signalline graph

Wavelet-Induced Rotary Encodings: RoPE Meets Graphs

Sep 26, 2025
IR
Isaac Reid
🏛️ University of Cambridge | Independent Researcher | Max Planck Institute for Intelligent Systems | University of Oxford | Alan Turing Institute | Google DeepMind

This work addresses the absence of rotation-equivariant positional encodings—specifically Rotary Position Embeddings (RoPE)—for non-grid graph-structured data. We propose WIRE, the first general-purpose RoPE extension to arbitrary graphs. WIRE constructs rotational transformations grounded in graph wavelet analysis, inherently satisfying permutation equivariance over node orderings and compatibility with linear attention mechanisms. Under mild conditions, it asymptotically approximates graph resistance distance, thereby explicitly encoding structural similarity. Parameter-free and plug-and-play, WIRE integrates seamlessly into diverse graph neural networks without architectural modification. Extensive experiments demonstrate that WIRE consistently outperforms existing positional encoding methods on tasks including monochromatic subgraph detection, point cloud semantic segmentation, and multiple standard graph benchmarks. Gains are especially pronounced on structure-sensitive tasks, validating both its theoretical foundation and empirical efficacy.

Extends Rotary Position Encodings to graph dataHandles node permutation equivariance and linear attentionImproves performance on graph-dependent tasks

Joint Generative Modeling of Scene Graphs and Images via Diffusion Models

Jan 02, 2024
BX
Bicheng Xu
🏛️ University of British Columbia | Vector Institute for AI

This paper introduces the first unconditional joint generation task of scene graphs and corresponding images, aiming to simultaneously synthesize structured scene graphs—comprising object categories, bounding boxes, and relational triplets—and photorealistic images from noise, enabling controllable and interpretable visual content generation. To this end, we propose DiffuseSG: a graph Transformer-based diffusion denoiser that unifies modeling of nodes (categories + coordinates), edges (relations), and adjacency matrices. We introduce IoU regularization and a continuous–discrete co-optimization mechanism, and pioneer the embedding of discrete category labels into a continuous latent space for joint diffusion modeling. Evaluated on Visual Genome and COCO-Stuff, DiffuseSG significantly outperforms state-of-the-art methods in both joint generation quality and fidelity. Moreover, it improves downstream scene graph completion and object detection performance, and generates high-fidelity samples that enhance model training through data augmentation.

Generating grounded scene graphs from noise for interpretable controlImproving performance in scene graph generation and downstream tasksModeling heterogeneous node and edge attributes in scene graphs

Latest Papers

What's happening recently
View more

This work addresses the challenge faced by visually impaired users in accessing node-link diagrams commonly distributed as bitmap images, a task for which existing assistive technologies are ill-suited due to their reliance on structured data rather than visual input. The paper presents the first lightweight deep learning approach for semantic segmentation of such diagram images, training a compact model on a large-scale synthetic dataset to achieve pixel-level parsing. The proposed method attains over 93% pixel accuracy on synthetic data and demonstrates strong performance both quantitatively and qualitatively. By enabling precise extraction of diagram semantics directly from rasterized images, this approach establishes a viable foundation for non-visual interaction and effectively bridges a critical gap in accessibility technology for bitmap-based graphical content.

accessibilityassistive technologybitmap images

This work addresses the challenge of efficient lossless compression for large-scale real-world graph data by proposing a novel algorithm that leverages geometric representations of graph structure through direct application of modern hyperbolic space embeddings. By capitalizing on the intrinsic hyperbolic geometry inherent in complex networks, the method achieves substantially improved compression efficiency while preserving lossless reconstruction. Experimental evaluation across diverse real-world graph datasets demonstrates that the proposed approach outperforms the current state-of-the-art methods by up to 42% in compression ratio, thereby validating the efficacy and superiority of hyperbolic embeddings for graph compression tasks.

graph compressionhyperbolic embeddingslossless compression

This work addresses the challenge of modeling component assembly relationships in 2D scenes under few-shot settings without semantic labels. The authors propose an end-to-end scene graph generation method that integrates geometric feature extraction with structural reasoning. Specifically, Faster R-CNN is employed to extract geometric representations of components, and a Transformer architecture constructs an initial adjacency matrix. To refine relational inference, the approach incorporates an attention-based Graph Convolutional Network (aGCN) with a message-passing mechanism. Notably, the model operates without semantic supervision and achieves accurate recovery of ground-truth assembly relationships using only a minimal number of training samples. Experimental results on a toy vehicle dataset demonstrate the method’s effectiveness, significantly advancing the capability to model assembly relationships in semantically unlabeled scenarios.

Assembly RelationshipComponent AssemblyGeometric Representation

Existing scene graph methods struggle to explicitly model the hierarchical entailment relations between locations and objects in Euclidean space, leading to insufficient structural consistency. This work proposes the first approach that incorporates hyperbolic geometry into scene graph representation learning, leveraging its innate capacity for encoding hierarchical structures. By integrating contrastive learning with attention mechanisms, the method enables more structured visual understanding. The resulting embeddings exhibit significantly improved hierarchical organization and semantic coherence, achieving a Graph IoU score of 33.51—representing an absolute gain of 8.14 over the strongest baseline—while maintaining competitive performance in retrieval tasks.

Euclidean geometryhierarchical relationshipshyperbolic space

Hot Scholars

HM

Huadong Ma

BUPT
Internet of ThingsMultimedia
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
YL

Yuxuan Liang

Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data MiningUrban ComputingUrban AIFoundation Models
PT

Ping Tan

Hong Kong University of Science and Technology (HKUST)
Computer VisionComputer Graphics
CL

Changsheng Lv

Beijing University of Posts and Telecommunications
Scene Graph GenerationAutonomous Driving