sequence-based geometry encoding

Designs, builds, or evaluates encodings that convert geometric structures and their connectivity into sequence or token representations suitable for sequence models; this includes coordinate-, delta-, and one-hot schemes, graph-to-sequence mappings (including Hamiltonian and hierarchical traversals), multiscale/multimodal fusion, and memory-efficient or hybrid designs. These encodings aim to preserve precise geometry and structural connectivity while enabling sequence-to-sequence tasks, encoding model fitting, and structured representation learning.

sequence-basedgeometryencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$179K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Representation of the structure of graphs by sequences of instructions

Dec 11, 2025
EL
Ezequiel López-Rubio
🏛️ University of Málaga

Existing graph representations—such as adjacency matrices—are incompatible with the text-processing paradigm of large language models (LLMs). To address this, we propose a reversible, locally structure-preserving mapping from graphs to instruction sequences: graphs are encoded into compact, deterministic instruction strings via adjacency matrix decomposition, yielding a parseable textual representation; a custom reversible parser enables lossless reconstruction. This work establishes the first bridge between graph algebra and LLM-native textual processing, simultaneously ensuring structural fidelity and sequence conciseness. Experiments demonstrate that our representation significantly improves LLM performance on graph modeling tasks—including graph classification and link prediction—validating both its effectiveness and generalizability across diverse graph domains.

Creates reversible, compact strings that preserve local graph patternsDevelops a graph representation method using instruction sequencesEnables deep learning models to process graph structures effectively

This work addresses the limitations of existing text-to-CAD generation methods, which often neglect assembly hierarchies and geometric constraints, leading to an excessively large search space and error accumulation. To overcome these challenges, the authors propose a hierarchical, geometry-aware graph representation that models parts and subassemblies as nodes and encodes geometric constraints as edges. The framework first predicts the structural layout and associated constraints, then leverages this information to guide the generation of CAD modeling operations and code. A novel structure-aware progressive curriculum learning strategy is introduced, employing controlled editing to construct tiered tasks and synthesize boundary cases. Additionally, the study presents the first Text-to-CAD dataset annotated with exploded views and explicit geometric constraints, along with tailored evaluation metrics. Experiments demonstrate that the proposed method significantly outperforms existing approaches in geometric fidelity and constraint satisfaction, with validation on a newly curated dataset of 12K samples.

geometric constraintsgeometry-awaregraph representation

This work proposes XShapeEnc, a training-free, general-purpose 2D shape encoding method that overcomes the limitations of conventional 1D positional encodings. By normalizing arbitrary 2D geometric shapes onto the unit disk, XShapeEnc represents geometry using orthogonal Zernike bases—either jointly or independently—and captures pose through harmonic orientation fields, while incorporating a frequency propagation mechanism to enhance high-frequency details. The method exhibits five desirable properties: invertibility, adaptability, spectral richness, robustness, and computational efficiency. Extensive experiments across diverse shape-aware tasks and on the newly introduced XShapeCorpus dataset demonstrate its theoretical soundness, computational efficiency, strong discriminative power, and broad applicability.

2D shape representationgeometric shapepositional encoding

Learning Mappings in Mesh-based Simulations

Jun 14, 2025
SH
Shirin Hosseinmardi
🏛️ University of California, Irvine

Modeling node-wise mappings on irregular point clouds over complex geometric domains remains challenging due to the absence of inherent structural regularity. To address this, we propose a parameter-free grid encoding method that projects point clouds onto regular lattices, yielding structured representations amenable to efficient end-to-end mapping learning and enabling full-response reconstruction from partial observations. Our approach integrates grid-footprint aggregation encoding, a lightweight E-UNet architecture, and an FFT-enhanced module. Evaluated across diverse 2D/3D physical simulation tasks, it achieves significant improvements in prediction accuracy, data efficiency, and robustness to noise. Compared with Fourier neural operators and Transformer-based baselines, our framework is notably lightweight, broadly generalizable, and deployment-friendly. It establishes a scalable, grid-aware modeling paradigm for computational science—bridging unstructured geometry and structured computation without architectural or parametric constraints.

Handling irregular mesh point clouds with limited tractabilityLearning mappings in geometrically complex mesh domainsRecovering full point cloud responses from partial observations

Fractal Language Modelling by Universal Sequence Maps (USM)

Aug 08, 2025
JS
Jonas S Almeida
🏛️ National Cancer Institute | INESC-ID | Instituto Superior Técnico | Universidade de Lisboa | IDMEC | Department of Biomedical Informatics | Stony Brook Medicine | Institute of Health Computing | University of Maryland | School of Medicine | Instituto de Engenharia de Sistemas e Computadores | University of Lisbon

This work addresses the challenge of numerically preserving contextual information across multiple scales and embedding dimensions in symbolic sequence modeling. We propose a bijective fractal encoding method based on Universal Sequence Mapping (USM), which employs bidirectional Chaos Game Representation (CGR) to achieve unbiased, invertible mapping from sequences to real-valued coordinates—eliminating seed-dependent bias in iterative constructions and ensuring strict one-to-one correspondence between numerical coordinates and sequence identities. Furthermore, we integrate Frequency-domain CGR (FCGR) with Chebyshev distance to enable non-integer *k*-mer frequency computation and multi-scale feature extraction without recomputation. Theoretically, we prove that USM converges to a stable embedding, thereby extending the theoretical foundations of fractal encoding. Experiments demonstrate the method’s universality and efficiency on both four-letter alphabets (e.g., genomic sequences) and arbitrary-size alphabets.

Develops fractal encoding for symbolic sequences at multiple scalesEnables efficient numeric embedding for arbitrary alphabet sizesResolves seeding biases in Universal Sequence Maps (USM)

Latest Papers

What's happening recently
View more

This study addresses the challenge that scientific foundation models, when applied to biological and physical systems, suffer from degraded geometric fidelity due to their discrete tokenized representations, which fail to preserve the intrinsic continuous geometric structure of such systems. By systematically comparing discrete-token and continuous-output heads under an identical encoder architecture, the work reveals the detrimental mechanism of the discrete bottleneck on geometric preservation. It introduces the novel concept of “geometric alignment tax” to explain how discretization induces geometric distortion and identifies three distinct representation failure modes. Through ablation studies on synthetic dynamical systems, rate–distortion theory, and MINE-based mutual information estimation, the authors demonstrate that while architectural performance gaps amount to only a 1.3× difference under continuous targets, they escalate dramatically to 3000× after discretization; moreover, no tested model simultaneously achieves low distortion, high mutual information, and global consistency.

continuous geometrydiscrete tokenizationfoundation models

This work addresses the challenge of efficient lossless compression for large-scale real-world graph data by proposing a novel algorithm that leverages geometric representations of graph structure through direct application of modern hyperbolic space embeddings. By capitalizing on the intrinsic hyperbolic geometry inherent in complex networks, the method achieves substantially improved compression efficiency while preserving lossless reconstruction. Experimental evaluation across diverse real-world graph datasets demonstrates that the proposed approach outperforms the current state-of-the-art methods by up to 42% in compression ratio, thereby validating the efficacy and superiority of hyperbolic embeddings for graph compression tasks.

graph compressionhyperbolic embeddingslossless compression

This work addresses the lack of theoretical foundations for substructure transferability in graph data by bridging transferable substructures with the intrinsic geometry of graph representation spaces from a functional behavior perspective. It proposes the first Riemannian geometry–based framework for learning intrinsic graph geometry, innovatively introducing neural vector bundles and local coordinate charts to construct the GAUGE pretraining architecture. A Dirichlet loss function is designed to enable explicit modeling of intrinsic graph geometry and quantification of transfer difficulty. The method demonstrates significant performance gains over existing models on zero-shot link prediction and graph isomorphism tasks, validating its expressive power and cross-task transferability.

graph foundation modelintrinsic geometryRiemannian geometry

This work addresses the limitations of existing representation alignment methods, which predominantly rely on geometric properties and struggle to capture the global structural organization of model representations. To overcome this, the study introduces topological data analysis into the field for the first time, proposing a Mapper-based visual analytics framework. By integrating force-directed layout, Bubble Sets, motif querying, and membrane-inspired heuristics, the framework enables a unified analytical pipeline spanning global structure alignment, local region matching, and fine-grained pattern exploration. Case studies on language and multimodal models, complemented by expert evaluations, demonstrate that the approach effectively reveals and compares the topological organization of representations across different models or layers, offering deep structural insights.

global structuremodel comparisonneural representations

Hot Scholars

TG

Travis Gagie

Associate Professor at Dalhousie University
data structuresdata compression
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
GG

Gennian Ge

Capital Normal University
CombinatoricsCoding theoryInformation Security
ZC

Zhengxue Cheng

Assistant Researcher, Shanghai Jiao Tong University
Video and Image CodingComputer VisionImage Quality Assessment
HH

Heng Huang

Brendan Iribe Endowed Professor in Computer Science, University Maryland College Park
Machine LearningAIBiomedical Data ScienceComputer Vision