Fractal Language Modelling by Universal Sequence Maps (USM)

📅 2025-08-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of numerically preserving contextual information across multiple scales and embedding dimensions in symbolic sequence modeling. We propose a bijective fractal encoding method based on Universal Sequence Mapping (USM), which employs bidirectional Chaos Game Representation (CGR) to achieve unbiased, invertible mapping from sequences to real-valued coordinates—eliminating seed-dependent bias in iterative constructions and ensuring strict one-to-one correspondence between numerical coordinates and sequence identities. Furthermore, we integrate Frequency-domain CGR (FCGR) with Chebyshev distance to enable non-integer *k*-mer frequency computation and multi-scale feature extraction without recomputation. Theoretically, we prove that USM converges to a stable embedding, thereby extending the theoretical foundations of fractal encoding. Experiments demonstrate the method’s universality and efficiency on both four-letter alphabets (e.g., genomic sequences) and arbitrary-size alphabets.

Technology Category

Cognitive Modeling & Cognitive Systems: Symbolic RepresentationsMachine Learning: Representation LearningData Mining & Knowledge Management: Data Compression

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Models for Web evolution
📝 Abstract
Motivation: With the advent of Language Models using Transformers, popularized by ChatGPT, there is a renewed interest in exploring encoding procedures that numerically represent symbolic sequences at multiple scales and embedding dimensions. The challenge that encoding addresses is the need for mechanisms that uniquely retain contextual information about the succession of individual symbols, which can then be modeled by nonlinear formulations such as neural networks. Context: Universal Sequence Maps(USM) are iterated functions that bijectively encode symbolic sequences onto embedded numerical spaces. USM is composed of two Chaos Game Representations (CGR), iterated forwardly and backwardly, that can be projected into the frequency domain (FCGR). The corresponding USM coordinates can be used to compute a Chebyshev distance metric as well as k-mer frequencies, without having to recompute the embedded numeric coordinates, and, paradoxically, allowing for non-integers values of k. Results: This report advances the bijective fractal encoding by Universal Sequence Maps (USM) by resolving seeding biases affecting the iterated process. The resolution had two results, the first expected, the second an intriguing outcome: 1) full reconciliation of numeric positioning with sequence identity; and 2) uncovering the nature of USM as an efficient numeric process converging towards a steady state sequence embedding solution. We illustrate these results for genomic sequences because of the convenience of a planar representation defined by an alphabet with only 4 tokens (the 4 nucleotides). Nevertheless, the application to alphabet of arbitrary cardinality was found to be straightforward.
Problem

Research questions and friction points this paper is trying to address.

Develops fractal encoding for symbolic sequences at multiple scales
Resolves seeding biases in Universal Sequence Maps (USM)
Enables efficient numeric embedding for arbitrary alphabet sizes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bijective fractal encoding with Universal Sequence Maps
Forward and backward Chaos Game Representations
Chebyshev distance and k-mer frequency computation
🔎 Similar Papers
No similar papers found.
Jonas S Almeida
Jonas S Almeida
Division of Cancer Epidemiology and Genetics
Data Science&Engineering
D
Daniel E Russ
National Cancer Institute, Division of Cancer Epidemiology and Genetics (DCEG), NCI Shady Grove, Rockville, MD 20850, USA.
Susana Vinga
Susana Vinga
INESC-ID, Instituto Superior Técnico, Universidade de Lisboa
BioinformaticsComputational BiologySystems BiomedicineMedical InformaticsAlignment-free sequence analysis
I
Ines Duarte
National Cancer Institute, Division of Cancer Epidemiology and Genetics (DCEG), NCI Shady Grove, Rockville, MD 20850, USA. INESC-ID, Instituto Superior Técnico, Universidade de Lisboa, 1000-029 Lisbon, Portugal.
Lee Mason
Lee Mason
National Cancer Institute, Division of Cancer Epidemiology and Genetics (DCEG), NCI Shady Grove, Rockville, MD 20850, USA.
P
Prahulla Bhawsar
National Cancer Institute, Division of Cancer Epidemiology and Genetics (DCEG), NCI Shady Grove, Rockville, MD 20850, USA. Department of Biomedical Informatics, Stony Brook Medicine, 100 Nicolls Rd, Stony Brook, NY 11794, USA.
A
Aaron Ge
Institute of Health Computing at the University of Maryland, School of Medicine, Baltimore MD 21201, USA.
A
Arlindo Oliveira
Instituto de Engenharia de Sistemas e Computadores, Instituto Superior Técnico, University of Lisbon, Lisbon 1000-029, Portugal.
J
Jeya B Balasubramanian
National Cancer Institute, Division of Cancer Epidemiology and Genetics (DCEG), NCI Shady Grove, Rockville, MD 20850, USA.