hamming lsh

Designs and implements locality-sensitive hashing functions and bucketed indexes that map fixed-length binary or discrete vectors to compact codes so items with small Hamming distance collide with high probability, enabling efficient candidate filtering. Builds and tunes hash bit selections, numbers of tables, and bucket/data structures and analyzes collision probabilities, sensitivity parameters, and query/space/time trade-offs for target Hamming-distance thresholds.

hamminglsh

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Faster and Space Efficient Indexing for Locality Sensitive Hashing

Mar 09, 2025
BD
Bhisham Dev Verma
🏛️ Wake Forest University | Indian Institute of Technology Hyderabad

To address the high computational complexity (O(md)) and large memory overhead of traditional locality-sensitive hashing (LSH) schemes—such as ELSH and SRP—in large-scale, high-dimensional approximate nearest neighbor search, this paper pioneers the integration of Count Sketch and its higher-order variants into LSH hash construction, proposing two novel LSH algorithms. Theoretically, our methods reduce hash computation complexity to O(d) and achieve space complexities of O(d) and O(N·d^{1/N}), respectively, while providing rigorous error bounds. Extensive experiments on multiple real-world datasets demonstrate that the proposed algorithms significantly accelerate hash construction, drastically reduce memory consumption, and maintain retrieval accuracy comparable to classical LSH baselines.

Improves time and space efficiency in LSH index construction.Introduces new algorithms for Euclidean distance and cosine similarity.Reduces hashcode computation complexity from O(md) to O(d).

This work addresses the limitations of binary locality-sensitive hashing (LSH) in approximate nearest neighbor (ANN) search, where recall and efficiency are often suboptimal. To overcome this, the authors propose a dynamic query modification mechanism that adaptively transforms the original query into a new center point at query time, significantly increasing both the probability and stability of hash collisions with true neighbors. Building upon this mechanism, they design MQ-Forest, an ANN retrieval framework that integrates random projection techniques for enhanced efficiency. Extensive experiments demonstrate that MQ-Forest reduces indexing and query time by up to 40% compared to baseline methods across multiple large-scale, high-dimensional datasets. Notably, this is the first approach to incorporate dynamic query transformation into binary LSH, effectively balancing accuracy and computational efficiency.

Approximate Near NeighbourBinary Locality Sensitive HashingHash Collision

PHast -- Perfect Hashing with fast evaluation

Apr 24, 2025
PB
Piotr Beling
🏛️ University of Łódź | Karlsruhe Institute of Technology

In static-set scenarios—such as database indexing and bioinformatics—existing perfect hash function (PHF) constructions struggle to simultaneously achieve high query speed and low construction overhead. To address this, we propose PHast, a novel PHF framework. Its key contributions are: (1) a synergistic design of compact primary hashing and linear mapping to minimize collision probability; (2) fixed-width seed encoding with overlapping value slicing and bumping to resolve seed collisions efficiently; and (3) a fully pipelined parallel construction algorithm that significantly improves scalability. Experiments on mainstream datasets show that PHast achieves the fastest known query performance—averaging under 1.1 CPU cycles per query—while reducing construction time by 1.3×–2.8× compared to state-of-the-art methods. Its space consumption remains comparable. Overall, PHast delivers superior end-to-end performance in both speed and efficiency.

Develops fast perfect hashing with minimal bits per keyEnables efficient parallel construction and compact seed encodingOptimizes bucket-placement for collision-free secondary hashing

Improving LSH via Tensorized Random Projection

Feb 11, 2024
BD
Bhisham Dev Verma
🏛️ Indian Institute of Technology Mandi | Indian Institute of Technology Hyderabad

This work addresses the exponential space complexity of conventional locality-sensitive hashing (LSH) in high-order tensor approximate nearest neighbor search, caused by explicit vectorization. We propose a novel LSH framework leveraging CP and tensor train (TT) decompositions—marking the first integration of low-rank tensor decomposition into LSH design. By operating directly on structured tensor representations, our method avoids vectorization entirely, reducing hash function parameter size from exponential to polynomial in tensor order while preserving sensitivity to both Euclidean distance and cosine similarity. We provide theoretical proof that the proposed scheme satisfies the formal LSH definition and offers probabilistic guarantees for approximate nearest neighbor retrieval. Empirical evaluation demonstrates substantial reductions in storage overhead, enables efficient low-rank tensor hashing, and confirms strong scalability. The approach thus bridges theoretical rigor with practical deployability for large-scale tensor similarity search.

Addressing exponential parameter growth in LSH.Improving LSH for tensor data efficiency.Proposing space-efficient LSH for Euclidean and cosine similarity.

Hierarchical Locality Sensitive Hashing for Structured Data: A Survey

Apr 24, 2022
WW
Wei Wu
🏛️ Central South University | Fudan University

Traditional locality-sensitive hashing (LSH) struggles to preserve topological and hierarchical relationships among elements in structured data—such as sequences, trees, and graphs—leading to inaccurate similarity estimation. To address this, this paper presents a systematic survey of hierarchical LSH (HLH) for structured data. We unify its development trajectory along three dimensions: data structures, application scenarios, and open challenges—offering the first such comprehensive characterization. Methodologically, we identify four core technical paradigms: (i) multi-granularity encoding, (ii) hierarchical similarity propagation, (iii) structure-aware signature generation, and (iv) recursive hashing via graph/tree decomposition. Based on these, we establish a taxonomy encompassing over ten state-of-the-art HLH algorithms, precisely delineating their applicability boundaries. Our work provides both a theoretical framework and practical guidelines for efficient approximate similarity search over structured data.

Efficient similarity computation for structured dataPreserving structural information in data similarityReviewing hierarchical LSH algorithms and applications

Latest Papers

What's happening recently
View more

This work investigates the approximability of Ulam and Cayley similarities within the locality-sensitive hashing (LSH) framework, quantifying their multiplicative distortion relative to similarity functions that admit exact LSH constructions. By integrating probabilistic analysis, combinatorics, and LSH theory, the study establishes the first sublinear upper bound of $O(n/\sqrt{\log n})$ and a lower bound of $\Omega(n^{0.12})$ on the LSH distortion for Ulam similarity. In contrast, it proves that the LSH distortion for Cayley similarity is tightly $\Theta(n)$. These results demonstrate that Ulam similarity admits efficient approximate nearest neighbor search with sublinear distortion, whereas Cayley similarity is fundamentally incompatible with LSH-based acceleration, thereby providing crucial theoretical foundations for permutation-based similarity search.

Cayley similaritylocality-sensitive hashingLSH distortion

Existing hash tables struggle to simultaneously achieve high operational speed and space efficiency. This work proposes Tiny Pointer Hash Tables (TPHT), the first practical hash table design that engineeringly integrates succinct pointers and compact key encoding, supporting dynamic resizing without global pauses and fully accommodating 64-bit keys. TPHT combines pointer compression, compact key representation, and two data layouts—chained and flattened—to ensure cache-friendly memory access and constant-time operations. Experimental results demonstrate that Chained-TPHT attains a space efficiency of 105.4%, while Flattened-TPHT achieves up to an 89.3% throughput improvement at 83.4% space efficiency, substantially advancing the latency–space Pareto frontier for hash table implementations.

hash tablesmemory overheadperformance

This study addresses the joint optimization of probe complexity and locality—defined as the maximum geometric distance between accessed memory cells—in open-addressing hash tables. Recognizing the inherent trade-off between linear probing and uniform probing, which struggle to balance both objectives, we first establish a theoretical lower bound of $\Omega(1/\varepsilon^2)$ on locality for any load factor $1-\varepsilon$. We then propose two algorithms: one achieving near-optimal $\tilde{O}(1/\varepsilon)$ bounds on both probe count and locality when the target load is known, and another greedy strategy that operates without prior knowledge of the load. Additionally, we prove that fixed-offset probing sequences yield an expected probe complexity of $O(\log n / \varepsilon^2)$ under arbitrary loads. Our analysis combines probabilistic methods, variance-based density bounds, and amortized techniques to establish matching upper and lower bounds for this problem.

hash tablesload factorlocality

This work addresses the lack of systematic research on non-minimal $k$-perfect hash functions ($k$-PHFs), which has hindered their adoption in high-speed static hash tables. We establish, for the first time, a tight space lower bound for any combination of load factor $\alpha \in (0,1]$ and $k \geq 1$, revealing that $\alpha < 1$ together with $k \geq 2$ can substantially reduce space overhead. Building upon PtrHash, we construct a practical $k$-PHF and integrate it into an efficient static hash table design featuring cache-line alignment. Experimental results demonstrate that our approach incurs only about 1.5× the theoretical space lower bound and achieves up to 1.5× faster negative and mixed query performance for $n \geq 30$ million keys, consistently outperforming existing hash set implementations.

cache efficiencyload factornon-minimal k-perfect hashing

This work addresses the inefficiency of insertions in bucketed cuckoo hashing under high load factors by proposing a novel insertion algorithm equipped with a prediction mechanism. By employing an improved random walk strategy, the method achieves—for the first time—a polynomial expected insertion time of $O(\delta^{-1}(\varepsilon^*)^{-1})$ with respect to $(\varepsilon^*)^{-1}$ when the load factor approaches the theoretical optimum $1 - (1+\delta)\varepsilon^*$, where $\delta \in [0.99^\ell, 1]$. It also guarantees $O(1)$ amortized expected eviction cost. Furthermore, during queries, the algorithm predicts the correct bucket for an element with probability $1 - o(1)$, enabling successful lookups with only $1 + o(1)$ bucket accesses and thereby significantly enhancing query performance.

amortized evictionsbucketized cuckoo hashinghash table

Hot Scholars

BY

Bian Yang

NTNU
biometricsidentity managementdata privacy and securitymultimedia coding and protection