subset pattern matching

Design, build, or analyze algorithms and data structures that perform exact subset-pattern matching and associative retrieval and that return exact nearest neighbors for discrete set or binary-vector representations. Work includes creating indexes, topology- or hash-based mappings, and query/update procedures with provable time and space bounds (e.g., constant-time associative lookup or exact Hamming-distance nearest-neighbor search).

subsetpatternmatching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.87
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Better Indexing for Rectangular Pattern Matching

Aug 24, 2025
PG
Paweł Gawrychowski
🏛️ University of Wrocław

This paper addresses the exact matching problem of arbitrary rectangular patterns in two-dimensional strings. Traditional indexing methods struggle to simultaneously support arbitrary rectangle shapes and efficient query processing. To overcome this, we propose the first index structure enabling near-linear query time, achieved by a divide-and-conquer strategy that maps the 2D matching problem to a series of 1D range queries via geometric interval encoding. Our construction integrates suffix arrays with striped range trees for precise pattern localization. The index occupies only $O(n log n)$ space, is built in $ ilde{O}(n)$ time, and answers queries for an $m$-character rectangular pattern in $O(m + k log^varepsilon n)$ time, where $k$ denotes the number of occurrences. Crucially, this is the first result to break prior conjectured lower bounds without assuming square-pattern restrictions—thereby advancing both the theoretical foundations and practical feasibility of two-dimensional pattern matching.

Achieving near-linear space and construction timeEfficient indexing for rectangular pattern matchingOvercoming lower bounds for query time complexity

Let them have CAKES: A Cutting-Edge Algorithm for Scalable, Efficient, and Exact Search on Big Data

Sep 11, 2023
ME
Morgan E. Prior
🏛️ Tufts University | University of Rhode Island

Exact k-NN search in large-scale, high-dimensional data suffers from poor efficiency, limited scalability, and difficulty adapting to arbitrary—especially non-metric—distance functions. Method: We propose three generic, exact k-NN algorithms grounded in metric geometry and the manifold hypothesis, introducing novel pruning and candidate filtering mechanisms. Their time complexity depends on the metric entropy and fractal dimension of the data—not on its cardinality or ambient dimension—enabling principled support for non-Euclidean distances such as Levenshtein and DTW. Contribution/Results: Implemented in Rust for high performance, our approach achieves >10× faster indexing over state-of-the-art methods on ANN-Benchmarks, genomic, and RF datasets; guarantees 100% recall in metric spaces; significantly outperforms SOTA in non-metric spaces; and exhibits near-constant scaling with data size.

Big DataEfficient AlgorithmK-Nearest Neighbor Search

This work addresses the memory bandwidth bottleneck in high-dimensional approximate nearest neighbor search (ANNS) on CPUs and GPUs, where conventional early termination mechanisms struggle to accelerate computation due to slow distance convergence. The authors propose a hardware-software co-design that integrates DIMM-level near-data processing (NDP) with a PCA-statistics-based feature-level early stopping mechanism, employing an estimate-and-correct strategy to accurately approximate full-dimensional distances. Additionally, they introduce bit-level dynamic floating-point compression and data-aware neighbor list mapping to substantially reduce memory access and communication overhead. Evaluated under strict accuracy constraints, the proposed system achieves 8.4× and 1.4× speedups over state-of-the-art CPU and GPU baselines, respectively, and outperforms the latest NDP accelerator, ANSMET, by 1.69×.

Approximate Nearest Neighbor SearchEarly ExitingMemory-Bound Computation

This work addresses the limitations of binary locality-sensitive hashing (LSH) in approximate nearest neighbor (ANN) search, where recall and efficiency are often suboptimal. To overcome this, the authors propose a dynamic query modification mechanism that adaptively transforms the original query into a new center point at query time, significantly increasing both the probability and stability of hash collisions with true neighbors. Building upon this mechanism, they design MQ-Forest, an ANN retrieval framework that integrates random projection techniques for enhanced efficiency. Extensive experiments demonstrate that MQ-Forest reduces indexing and query time by up to 40% compared to baseline methods across multiple large-scale, high-dimensional datasets. Notably, this is the first approach to incorporate dynamic query transformation into binary LSH, effectively balancing accuracy and computational efficiency.

Approximate Near NeighbourBinary Locality Sensitive HashingHash Collision

A Comprehensive Survey on Vector Database: Storage and Retrieval Technique, Challenge

Oct 18, 2023
YH
Yikun Han
🏛️ University of Michigan | Chinese Academy of Sciences

Managing and retrieving high-dimensional vector data poses significant challenges, particularly as traditional databases fail to meet performance requirements and the need for tight integration with large language models (LLMs) intensifies. Method: This paper systematically surveys four major approximate nearest neighbor search (ANNS) paradigms—hashing, tree-based indexing, graph-based methods (e.g., HNSW), and quantization (PQ/SQ)—and integrates hybrid optimization strategies. Contribution/Results: It introduces, for the first time, a “Four-Dimensional Methodology” framework tailored for industrial deployment of vector databases, analyzing trade-offs among accuracy, latency, memory footprint, and scalability. The work constructs a structured knowledge graph covering 200+ ANNS algorithms and proposes a novel paradigm for deep synergy between vector databases and LLMs. Collectively, these contributions provide both theoretical foundations and practical guidelines for system selection, architectural design, and development of AI-native database systems.

Compare advanced VDB solutions with strengths and limitationsExplore coupling VDBs with large language modelsReview storage and retrieval techniques in vector databases

Latest Papers

What's happening recently
View more

This work addresses the design of nearest neighbor search data structures tailored to a given query distribution. It proposes the first algorithm capable of efficiently learning an approximately optimal balanced halfspace partitioning tree under Gaussian-like distributional assumptions. By formulating tree construction as a balanced halfspace cut problem and integrating polynomial threshold functions with a distribution-aware learning strategy, the method circumvents the NP-hard regularized optimization typically involved. Under the assumption that a perfect partitioning tree exists, the approach achieves query time better than $O(nd)$ while ensuring provably bounded cutting error in the learned tree, thereby significantly enhancing both the efficiency and theoretical guarantees of data-driven nearest neighbor search.

balanced halfspace treesdata-driven algorithm designnearest neighbor search

This work addresses the challenge of efficient approximate nearest neighbor search and maximum inner product search (MIPS) in high-dimensional embeddings from large language models. The authors propose a novel approach that transforms asymmetric MIPS into Euclidean nearest neighbor search via dimension augmentation, combined with Equi-Voronoi Polytopes (EVP) quantization and a Fast Linear Assignment Sorting (FLAS) one-dimensional pre-sorting mechanism. This integration substantially accelerates k-nearest neighbor graph (kNNG) construction and query processing while enhancing memory access locality and cache efficiency. Evaluated in the SISAP 2026 challenge, the method achieves low-latency, high-recall MIPS performance, significantly outperforming existing baselines.

Approximate Nearest Neighbor SearchHigh-Dimensional Embeddingsk-Nearest Neighbor Graph

This work addresses the high latency and service disruption caused by frequent index rebuilds in existing approximate nearest neighbor (ANN) methods under dynamic vector database updates. To overcome these limitations, the authors propose ACRONYM—a co-designed algorithm-hardware platform that leverages a data-distribution-agnostic XOR-and-Accumulate (XAC) systolic array encoder and Hamming-distance-based search, integrated with content-addressable memory (CAM) to enable in-memory parallel computation. A two-stage coarse-to-fine retrieval architecture circumvents CAM dimensionality constraints, allowing continuous, interruption-free updates. Evaluated on million-scale dynamic datasets, ACRONYM achieves over 90% recall, 8 million queries per second throughput, only 32 MB memory footprint, and 2.56 μJ per query energy efficiency—outperforming CPU-based HNSW by 400× and GPU-based FAISS-IVF by 80× in speed.

approximate nearest neighbor searchcontinuous updatesdynamic vector databases

This work addresses the challenge of efficiently supporting both vector similarity search and arbitrary attribute filtering in high-dimensional approximate nearest neighbor retrieval. The authors propose a lightweight graph-based indexing algorithm that seamlessly integrates attribute filtering into the graph traversal process, overcoming the efficiency bottlenecks of existing methods when handling unseen query vectors combined with complex attribute constraints. Experimental results on multiple real-world datasets demonstrate that the proposed approach significantly outperforms state-of-the-art techniques, achieving substantially faster query latency while maintaining high recall. The method thus offers a compelling balance among flexibility, efficiency, and scalability for hybrid vector-and-attribute search scenarios.

arbitrary attribute combinationsattribute filteringefficient search

This work addresses the problem of cardinality estimation for similarity queries in high-dimensional spaces by proposing a novel method that balances accuracy and online efficiency. The approach leverages locality-sensitive hashing (LSH) to partition the space and integrates adaptive multi-probe bucket probing, progressive sampling, and asymmetric distance computation. It also supports dynamic data updates, making it suitable for evolving datasets. Experimental results demonstrate that the proposed scheme significantly outperforms existing methods across multiple high-dimensional datasets, achieving high estimation accuracy while substantially improving online query response time. The method is thus well-suited for large-scale applications involving both static and dynamic data.

adaptive bucket probingcardinality estimationhigh-dimensional spaces

Hot Scholars

TS

Tatiana Starikovskaya

Ecole Normale Supérieure
Stringologyrandomized algorithmsapproximate algorithmsstreaming algorithms
JZ

Jingyu Zhang

WNLO Huazhong University of Science and Technology
optical
BV

Benjamin Van Durme

Johns Hopkins University / Microsoft
LinguisticsNatural Language ProcessingArtificial Intelligence
LW

Longyue Wang

Alibaba International
Large Language ModelMachine TranslationNatural Language ProcessingLanguange Agent
DK

Daniel Khashabi

Johns Hopkins University
Natural Language ProcessingArtificial IntelligenceMachine Learning