Score
Designs and evaluates hash functions and encoding schemes that map inputs into compact, similarity-preserving representations—including locality-sensitive and feature hashing, perceptual hashes, and similarity-preserving key-dependent encodings—so that similarity relationships and nearest-neighbor behavior are approximately retained for efficient search, storage, or reduced-cost model input. Also develops key-dependent and homomorphic-friendly hash constructions that enable secure or computation-preserving operations on hashed data and permit native AI/ML algorithms to run on compressed/hashed representations without modifying those algorithms, while managing collision, distortion, and approximation trade-offs.
High-dimensional vector approximate nearest neighbor search (ANNS) suffers from efficiency bottlenecks due to linear growth of distance computation cost with dimensionality—especially acute for LLM-derived semantic vectors. This work systematically evaluates six dimensionality reduction (DR) techniques—PCA, product quantization, autoencoders, contrastive learning-based DR, LSH variants, and random projection—quantifying their acceleration effects on mainstream ANNS engines (e.g., FAISS, Annoy) under a unified experimental framework. We propose two analyzable DR–search co-design architectures and theoretically derive critical pruning gain thresholds, characterizing the fundamental trade-off between dimensionality compression and retrieval accuracy degradation. Experiments across six public benchmarks show that deep DR methods achieve 1.8–3.5× speedup while maintaining >90% recall. Furthermore, we provide a data-aware guideline for optimal DR technique selection based on intrinsic data properties.
This work investigates the approximability of Ulam and Cayley similarities within the locality-sensitive hashing (LSH) framework, quantifying their multiplicative distortion relative to similarity functions that admit exact LSH constructions. By integrating probabilistic analysis, combinatorics, and LSH theory, the study establishes the first sublinear upper bound of $O(n/\sqrt{\log n})$ and a lower bound of $\Omega(n^{0.12})$ on the LSH distortion for Ulam similarity. In contrast, it proves that the LSH distortion for Cayley similarity is tightly $\Theta(n)$. These results demonstrate that Ulam similarity admits efficient approximate nearest neighbor search with sublinear distortion, whereas Cayley similarity is fundamentally incompatible with LSH-based acceleration, thereby providing crucial theoretical foundations for permutation-based similarity search.
This work addresses the limitations of binary locality-sensitive hashing (LSH) in approximate nearest neighbor (ANN) search, where recall and efficiency are often suboptimal. To overcome this, the authors propose a dynamic query modification mechanism that adaptively transforms the original query into a new center point at query time, significantly increasing both the probability and stability of hash collisions with true neighbors. Building upon this mechanism, they design MQ-Forest, an ANN retrieval framework that integrates random projection techniques for enhanced efficiency. Extensive experiments demonstrate that MQ-Forest reduces indexing and query time by up to 40% compared to baseline methods across multiple large-scale, high-dimensional datasets. Notably, this is the first approach to incorporate dynamic query transformation into binary LSH, effectively balancing accuracy and computational efficiency.
This paper addresses the problem of efficient privacy-preserving Hamming distance computation under Property-Preserving Hashing (PPH). We propose the first PPH scheme that enables constant-time approximate distance estimation without oracle access. Methodologically, we build upon a threshold evaluation framework, integrating binary search, constant-round query optimization, and a customized PPH construction—leveraging cryptographic hashing and probabilistic approximation—to achieve sublinear, even constant-time distance estimation using ciphertexts only. Theoretically, our scheme satisfies strong cryptographic assumptions (e.g., DDH). Empirically, it significantly outperforms existing baselines while maintaining high accuracy. Our core contribution is the first realization of both efficiency and rigorous privacy preservation under strict security constraints, thereby breaking the performance bottleneck in PPH-based similarity search.
Traditional locality-sensitive hashing (LSH) struggles to preserve topological and hierarchical relationships among elements in structured data—such as sequences, trees, and graphs—leading to inaccurate similarity estimation. To address this, this paper presents a systematic survey of hierarchical LSH (HLH) for structured data. We unify its development trajectory along three dimensions: data structures, application scenarios, and open challenges—offering the first such comprehensive characterization. Methodologically, we identify four core technical paradigms: (i) multi-granularity encoding, (ii) hierarchical similarity propagation, (iii) structure-aware signature generation, and (iv) recursive hashing via graph/tree decomposition. Based on these, we establish a taxonomy encompassing over ten state-of-the-art HLH algorithms, precisely delineating their applicability boundaries. Our work provides both a theoretical framework and practical guidelines for efficient approximate similarity search over structured data.
This work addresses the exponential space complexity of conventional locality-sensitive hashing (LSH) in high-order tensor approximate nearest neighbor search, caused by explicit vectorization. We propose a novel LSH framework leveraging CP and tensor train (TT) decompositions—marking the first integration of low-rank tensor decomposition into LSH design. By operating directly on structured tensor representations, our method avoids vectorization entirely, reducing hash function parameter size from exponential to polynomial in tensor order while preserving sensitivity to both Euclidean distance and cosine similarity. We provide theoretical proof that the proposed scheme satisfies the formal LSH definition and offers probabilistic guarantees for approximate nearest neighbor retrieval. Empirical evaluation demonstrates substantial reductions in storage overhead, enables efficient low-rank tensor hashing, and confirms strong scalability. The approach thus bridges theoretical rigor with practical deployability for large-scale tensor similarity search.
This work addresses the problem of angular approximate nearest neighbor (Angular ANN) search on high-dimensional spheres by proposing a unified framework that systematically integrates locality-sensitive hashing (LSH) and locality-sensitive filtering (LSF). By constructing a novel LSF-based data structure, the paper reformulates LSF theory in an expository “guided tour” manner, revealing its deep connections with LSH and strengthening key lemmas to establish the optimality of the proposed structure in terms of both query and space complexity. The study not only delivers a concise and rigorous theoretical analysis but also provides a clear, coherent entry point and survey perspective for future research on Angular ANN.
Existing deep hashing methods are limited by the representational discrepancy between continuous features and discrete Hamming codes. This work proposes HashViT, the first framework to natively embed hash learning within a Vision Transformer by introducing a dedicated HASH token that progressively evolves binary codes layer-by-layer inside the network, thereby circumventing end-of-pipeline quantization. The HASH token comprises a Hash Register and a Semantic Workspace, complemented by a lightweight Hash Refinement Adapter for fine-grained optimization. Through a joint training strategy integrating learnable semantic centroid supervision, class-token similarity distillation, and quantization regularization, HashViT achieves state-of-the-art or highly competitive performance on three mainstream image retrieval benchmarks while preserving the efficiency advantages of compact Hamming codes for fast retrieval.
This work addresses the absence of locality-sensitive hashing (LSH) schemes tailored for approximate nearest neighbor (ANN) search in hyperbolic space by proposing the first native LSH construction for this geometry. The method introduces a two-dimensional hashing scheme based on hyperbolic hyperplane rounding and extends it to higher dimensions via dimensionality reduction combined with local isometric embeddings. Theoretical analysis establishes an upper bound on the performance parameter ρ of ρ ≤ 1/c for d = 2 and ρ ≤ 1.59/c for d ≥ 3, along with a lower bound of ρ ≥ 1/c². This approach achieves, for the first time, sublinear query time and storage complexity for ANN search in hyperbolic space with rigorous theoretical guarantees.
This work addresses the limited representational capacity of existing methods in complex scenes by proposing a novel neural network architecture based on multi-scale feature fusion and an adaptive attention mechanism. By dynamically integrating local details with global semantic information, the proposed approach significantly enhances model robustness under challenging conditions such as occlusion, illumination variations, and background clutter. Extensive experiments demonstrate that the model achieves state-of-the-art performance across multiple benchmark datasets while maintaining low computational overhead, offering a practical solution for real-world deployment. The primary contribution lies in the design of a lightweight yet highly effective feature interaction mechanism, whose efficacy in improving generalization capability is systematically validated.
This work addresses the challenge in unsupervised fine-grained image hashing where neglecting collision resistance often causes semantically similar yet distinct samples to be mapped to identical hash codes. To this end, the paper proposes the CS3H framework, which explicitly models collision resistance for the first time in this task. Specifically, it introduces a normalized Hamming distance loss computed in a single forward pass to directly optimize similarity in Hamming space, and incorporates a collision-aware attention module to enhance the learning of rare yet discriminative local features. Extensive experiments demonstrate that CS3H significantly outperforms existing methods across multiple benchmark datasets, achieving higher retrieval accuracy and stronger collision resistance while maintaining low computational overhead.