lsh channel grouping

Designs, implements, and evaluates grouping methods that use locality-sensitive hashing to cluster similar channels (or per-channel/patch embeddings) by hashing their representations into buckets so similar items collide. Builds hash-based, training-free nearest-neighbor retrieval and grouping pipelines for operations such as channel merging, pruning, or patch-wise compression, and analyzes hashing parameters and grouping quality versus downstream efficiency and accuracy.

lshchannelgrouping

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the problem of cardinality estimation for similarity queries in high-dimensional spaces by proposing a novel method that balances accuracy and online efficiency. The approach leverages locality-sensitive hashing (LSH) to partition the space and integrates adaptive multi-probe bucket probing, progressive sampling, and asymmetric distance computation. It also supports dynamic data updates, making it suitable for evolving datasets. Experimental results demonstrate that the proposed scheme significantly outperforms existing methods across multiple high-dimensional datasets, achieving high estimation accuracy while substantially improving online query response time. The method is thus well-suited for large-scale applications involving both static and dynamic data.

adaptive bucket probingcardinality estimationhigh-dimensional spaces

This work addresses the problem of angular approximate nearest neighbor (Angular ANN) search on high-dimensional spheres by proposing a unified framework that systematically integrates locality-sensitive hashing (LSH) and locality-sensitive filtering (LSF). By constructing a novel LSF-based data structure, the paper reformulates LSF theory in an expository “guided tour” manner, revealing its deep connections with LSH and strengthening key lemmas to establish the optimality of the proposed structure in terms of both query and space complexity. The study not only delivers a concise and rigorous theoretical analysis but also provides a clear, coherent entry point and survey perspective for future research on Angular ANN.

angular distanceApproximate Near Neighborhigh-dimensional search

Faster and Space Efficient Indexing for Locality Sensitive Hashing

Mar 09, 2025
BD
Bhisham Dev Verma
🏛️ Wake Forest University | Indian Institute of Technology Hyderabad

To address the high computational complexity (O(md)) and large memory overhead of traditional locality-sensitive hashing (LSH) schemes—such as ELSH and SRP—in large-scale, high-dimensional approximate nearest neighbor search, this paper pioneers the integration of Count Sketch and its higher-order variants into LSH hash construction, proposing two novel LSH algorithms. Theoretically, our methods reduce hash computation complexity to O(d) and achieve space complexities of O(d) and O(N·d^{1/N}), respectively, while providing rigorous error bounds. Extensive experiments on multiple real-world datasets demonstrate that the proposed algorithms significantly accelerate hash construction, drastically reduce memory consumption, and maintain retrieval accuracy comparable to classical LSH baselines.

Improves time and space efficiency in LSH index construction.Introduces new algorithms for Euclidean distance and cosine similarity.Reduces hashcode computation complexity from O(md) to O(d).

Improving LSH via Tensorized Random Projection

Feb 11, 2024
BD
Bhisham Dev Verma
🏛️ Indian Institute of Technology Mandi | Indian Institute of Technology Hyderabad

This work addresses the exponential space complexity of conventional locality-sensitive hashing (LSH) in high-order tensor approximate nearest neighbor search, caused by explicit vectorization. We propose a novel LSH framework leveraging CP and tensor train (TT) decompositions—marking the first integration of low-rank tensor decomposition into LSH design. By operating directly on structured tensor representations, our method avoids vectorization entirely, reducing hash function parameter size from exponential to polynomial in tensor order while preserving sensitivity to both Euclidean distance and cosine similarity. We provide theoretical proof that the proposed scheme satisfies the formal LSH definition and offers probabilistic guarantees for approximate nearest neighbor retrieval. Empirical evaluation demonstrates substantial reductions in storage overhead, enables efficient low-rank tensor hashing, and confirms strong scalability. The approach thus bridges theoretical rigor with practical deployability for large-scale tensor similarity search.

Addressing exponential parameter growth in LSH.Improving LSH for tensor data efficiency.Proposing space-efficient LSH for Euclidean and cosine similarity.

Hashing for Structure-Based Anomaly Detection

May 16, 2025
FL
Filippo Leveni
🏛️ Politecnico di Milano | Università della Svizzera italiana

This work addresses anomaly detection on low-dimensional manifold-structured data. We propose an efficient isolation-based method that embeds data into a high-dimensional semantic-enhanced preference space and employs Locality-Sensitive Hashing (LSH) to accelerate sparse neighborhood estimation, thereby identifying the most isolated samples as anomalies. To our knowledge, this is the first approach to integrate LSH into a preference-space isolation framework, achieving both theoretical soundness and computational efficiency. Extensive experiments on multiple benchmark datasets demonstrate state-of-the-art detection performance, with inference speed improved by 3–5× over existing methods, alongside substantial reductions in time and memory overhead. The source code is publicly available.

Detecting isolated points in high-dimensional Preference SpaceIdentifying anomalies in structured low-dimensional manifoldsImproving efficiency with Locality Sensitive Hashing technique

Latest Papers

What's happening recently
View more

This work addresses the limitations of binary locality-sensitive hashing (LSH) in approximate nearest neighbor (ANN) search, where recall and efficiency are often suboptimal. To overcome this, the authors propose a dynamic query modification mechanism that adaptively transforms the original query into a new center point at query time, significantly increasing both the probability and stability of hash collisions with true neighbors. Building upon this mechanism, they design MQ-Forest, an ANN retrieval framework that integrates random projection techniques for enhanced efficiency. Extensive experiments demonstrate that MQ-Forest reduces indexing and query time by up to 40% compared to baseline methods across multiple large-scale, high-dimensional datasets. Notably, this is the first approach to incorporate dynamic query transformation into binary LSH, effectively balancing accuracy and computational efficiency.

Approximate Near NeighbourBinary Locality Sensitive HashingHash Collision

This work addresses the absence of locality-sensitive hashing (LSH) schemes tailored for approximate nearest neighbor (ANN) search in hyperbolic space by proposing the first native LSH construction for this geometry. The method introduces a two-dimensional hashing scheme based on hyperbolic hyperplane rounding and extends it to higher dimensions via dimensionality reduction combined with local isometric embeddings. Theoretical analysis establishes an upper bound on the performance parameter ρ of ρ ≤ 1/c for d = 2 and ρ ≤ 1.59/c for d ≥ 3, along with a lower bound of ρ ≥ 1/c². This approach achieves, for the first time, sublinear query time and storage complexity for ANN search in hyperbolic space with rigorous theoretical guarantees.

Approximate Nearest Neighbor SearchHyperbolic SpaceLocality Sensitive Hashing

This work addresses the high computational cost of clustering high-dimensional vectors, which poses a significant bottleneck for large-scale indexing in vector retrieval systems. To overcome this challenge, the authors propose SuperKMeans, a novel approach that integrates dimension pruning with a recall-based early stopping mechanism. This combination substantially accelerates k-means training while preserving the quality of cluster centroids. Empirical evaluations demonstrate that SuperKMeans achieves a 7× speedup over FAISS and Scikit-Learn on CPU and outperforms cuVS by 4× on GPU, all without compromising clustering accuracy. The method thus enables efficient, high-quality clustering of high-dimensional data, making it well-suited for deployment in large-scale vector search infrastructures.

clusteringhigh-dimensional datak-means

This work investigates the approximability of Ulam and Cayley similarities within the locality-sensitive hashing (LSH) framework, quantifying their multiplicative distortion relative to similarity functions that admit exact LSH constructions. By integrating probabilistic analysis, combinatorics, and LSH theory, the study establishes the first sublinear upper bound of $O(n/\sqrt{\log n})$ and a lower bound of $\Omega(n^{0.12})$ on the LSH distortion for Ulam similarity. In contrast, it proves that the LSH distortion for Cayley similarity is tightly $\Theta(n)$. These results demonstrate that Ulam similarity admits efficient approximate nearest neighbor search with sublinear distortion, whereas Cayley similarity is fundamentally incompatible with LSH-based acceleration, thereby providing crucial theoretical foundations for permutation-based similarity search.

Cayley similaritylocality-sensitive hashingLSH distortion

This work addresses a key limitation in existing deep semantic hashing methods, where fixed-width and fixed-position semantic channels induce discontinuities in the loss function, thereby hindering optimization. To overcome this, the authors propose the Dynamic Semantic Channel Hashing (DSCH) loss, which dynamically adjusts both the position and scale of semantic channels to yield a smoother loss landscape and enhance hash code learning. Additionally, they introduce a tie-aware mean Average Precision (mAP) metric to more accurately evaluate retrieval performance under discrete Hamming distances. Evaluated across both cross-modal and single-modal settings on two benchmark datasets using two distinct architectures, DSCH significantly outperforms current state-of-the-art methods in 35 out of 40 tasks, achieving up to a 1.75 percentage point improvement in tie-aware mAP.

cross-modal retrievalHamming spaceloss function

Hot Scholars

CY

Chen Yan

Associate Professor, Zhejiang University, College of EE
CPS SecurityEmbedded System SecuritySensor Security
ZA

Zhulin An

Institute Of Computing Technology Chinese Academy Of Sciences
Automatic Deep LearningLifelong Learning
RP

Rameshwar Pratap

IIT Hyderabad
Algorithms for big datasketching/sampling algorithmsmachine learningdata mining
XZ

Xiaofang Zhou

Hong Kong University of Science and Technology
databasesbig datadata scienceAI
RZ

Ruiyuan Zhang

Zhejiang University
MultiModal3D Part AssemblyMixture-of-Expert