Score
Design and implement compact representations that embed partial orders or hierarchical structures into vectors or indexable encodings that preserve the order relation, using techniques such as order-embeddings, nested-set interval encodings, and chain decompositions. Build algorithms to decompose low-width DAGs into chains and to support efficient subsumption/ancestor–descendant queries and ordering tests on the encoded structure.
This work addresses the problem of constructing low-dimensional geometric embeddings for directed acyclic graphs (DAGs) with ancestor–descendant relationships, aiming to avoid embedding dimensions that scale explosively with the number of nodes or graph depth. By leveraging structural properties—such as treewidth and the number of cross edges—and insights from geometric embedding theory, the authors propose a compact representation whose dimension depends only on these structural parameters rather than the total node count. Key contributions include a proof that any directed tree admits an exact reachability-preserving embedding in three dimensions, and an upper bound of \( O(t \log n) \) dimensions for DAGs of treewidth \( t \), accompanied by a nearly matching lower bound that reveals fundamental limits on dimensionality. Experiments on real-world datasets demonstrate that the method substantially reduces embedding dimension while maintaining high recall, outperforming existing approaches with theoretical guarantees.
Segmented linear approximation (PLA) in learned indexes suffers from suboptimal storage efficiency, and no information-theoretic space lower bound exists for PLA under both compression and indexing constraints. Method: We establish the first information-theoretic space lower bound for PLA in these dual settings, then design a novel, minimalist data structure that achieves theoretically optimal compact representation for 2D monotonic point sequences under a given error bound. The structure supports O(log n)-time x-value lookup and segment evaluation. Contribution/Results: Our approach unifies the modeling of PLA’s compressibility and queryability—yielding the first systematic lower-bound analysis, constructive guarantee, and efficient implementation for PLA-based learned indexes. The space usage is asymptotically tight to the lower bound, achieving succinctness on most practical distributions. This work bridges a critical theoretical and engineering gap in learned indexing research.
Standard $k^2$-trees suffer from redundant subtree storage, poor cache locality, and high navigation overhead due to breadth-first traversal when compressing sparse binary matrices. To address these issues, this paper introduces the first depth-first representation of the $k^2$-tree—termed DF-$k^2$-tree. Our method employs depth-first encoding, hash-based detection of subtree isomorphism, compact bitmap serialization, and cache-aware access patterns to achieve linear-time duplicate-subtree identification and superior compression ratios. Evaluated on web-graph adjacency matrices, DF-$k^2$-tree significantly outperforms the standard $k^2$-tree in compression ratio and accelerates sparse matrix multiplication across diverse real-world sparse matrices. The core innovation lies in shifting the traversal paradigm from breadth-first to depth-first—thereby simultaneously enhancing space efficiency, navigation speed, and cache locality—without compromising structural integrity or query support.
This work addresses the poor cache locality of the traditional $k^2$-tree, which employs a breadth-first layout and hinders the efficiency of matrix operations. The paper presents the first systematic formulation of depth-first $k^2$-tree representations, introducing EDF-1, BP, and their compressed variants CEDF and CBP. By leveraging balanced parentheses sequences (DFUDS), suffix arrays, and LCP arrays, the authors design linear-time algorithms to identify and compress repeated subtrees. This approach substantially improves memory locality and reduces peak memory consumption. Among the proposed variants, CEDF achieves the highest compression ratio, while both EDF-1 and CEDF demonstrate superior performance across diverse matrix operations and datasets.
Existing balanced tree structures—such as AVL trees—lack tight information-theoretic lower bounds on encoding size, suffer from intractable exact enumeration, and offer limited support for efficient static queries. Method: We introduce a novel tree decomposition framework grounded in generating functions and combinatorial enumeration, enabling rigorous asymptotic analysis; we further design a succinct data structure supporting constant-time queries—including ancestor, subtree size, and level-order traversal—and generalize our approach to recursively defined balanced tree families (e.g., red-black trees, WB-trees) via functional equations. Contribution/Results: We establish the first provable information-theoretic lower bound for AVL tree encoding—approximately 0.938 bits per node—and present a succinct representation achieving this bound while supporting rich navigational queries. Our unified framework extends to broader classes of height-balanced trees, bridging theoretical limits and practical succinct data structure design.
This study addresses how to preserve relational structure rather than individual element information under constrained representations. To this end, it constructs a unified relational compression framework that integrates graph summarization and spectral sparsification within a common interface. Furthermore, the work introduces a finite-codeword collision model, establishing precise correspondences among relational geometry, Rényi-2 occupancy, and spherical geometry. Adopting a “source–description–reconstruction” paradigm, the proposed approach synthesizes techniques from graph theory, spectral analysis, and relational distillation. Through evaluations on both graph and image tasks, the study demonstrates complementary pathways for diverse relational requirements, achieves a unified assessment of constrained representations, and provides a novel theoretical foundation for relational information compression.
This work addresses the challenge of efficiently and unambiguously encoding arbitrary finite simple graphs into compact strings to facilitate graph similarity computation, generation, and integration with language models. The authors propose IsalGraph, a method that constructs graphs within a virtual machine using a nine-character instruction set, leveraging a circular doubly linked list and a dual-pointer traversal mechanism to achieve a lossless mapping from graphs to valid strings. Greedy and backtracking algorithms are designed to produce lexicographically shortest canonical strings, guaranteeing graph isomorphism invariance and compatibility with language models. Experiments on five benchmarks—IAM Letter (LOW/MED/HIGH), LINUX, and AIDS—demonstrate that Levenshtein distances between IsalGraph strings strongly correlate with graph edit distances, confirming the approach’s effectiveness for graph similarity and generation tasks.
This work addresses the challenge of efficiently performing breadth-first search (BFS) and related graph algorithms directly on highly compressed representations of planar graphs while strictly limiting auxiliary space usage. The authors propose the first succinct encoding of planar graphs that can be constructed in linear time, supports native BFS traversal, and requires only o(n) additional bits of space. This representation enables constant-time queries on both the BFS tree and the original graph structure. Furthermore, it extends to several advanced applications, including computing balanced separators of size O(√n), diameter-based tree decompositions, triangulation, and bipartiteness testing. By integrating techniques from succinct data structures, planar graph embedding theory, and dual-graph traversal, the method achieves sublinear space complexity without sacrificing linear-time efficiency.
This work addresses the challenge of efficiently compressing multiple sets while supporting fast queries. The authors propose a compression scheme based on inter-set differences, leveraging a minimum spanning tree (MST) to optimize differential encoding across sets, thereby significantly improving compression efficiency. They further design a data structure that supports fundamental operations—including membership testing, predecessor/successor queries, and random access—all executed in logarithmic time. Experimental results demonstrate that the proposed method outperforms existing standard approaches in both construction speed and query performance, achieving an effective balance between space efficiency and time efficiency.
本文讨论了超树分解中的重根性问题,通过定义一种放宽的规范形式来解决该问题,从而获得可重根且易于处理的分解类。