Score
Design and implement hashing functions that map continuous latent representations to compact binary Hamming codes by anchoring learned prototypes in the latent space; build alignment mechanisms (global, stochastic, or contrastive stochastic neighborhood alignment) that transfer and preserve global neighborhood and semantic similarity structure into the binary space.
Existing deep hashing methods are limited by the representational discrepancy between continuous features and discrete Hamming codes. This work proposes HashViT, the first framework to natively embed hash learning within a Vision Transformer by introducing a dedicated HASH token that progressively evolves binary codes layer-by-layer inside the network, thereby circumventing end-of-pipeline quantization. The HASH token comprises a Hash Register and a Semantic Workspace, complemented by a lightweight Hash Refinement Adapter for fine-grained optimization. Through a joint training strategy integrating learnable semantic centroid supervision, class-token similarity distillation, and quantization regularization, HashViT achieves state-of-the-art or highly competitive performance on three mainstream image retrieval benchmarks while preserving the efficiency advantages of compact Hamming codes for fast retrieval.
Existing hash-center-based methods suffer from neglecting inter-class semantic relationships due to random initialization; while two-stage optimization alleviates this issue, it introduces error accumulation and computational redundancy. This paper proposes an end-to-end joint learning framework that dynamically reallocates hash centers within a predefined codebook to adaptively model inter-class semantic structure, while co-optimizing neural hash functions. Our key contributions are: (1) a dynamic center matching mechanism that eliminates the need for explicit center optimization stages; (2) multi-head representation fusion to enhance semantic discriminability; and (3) a codebook-based lightweight architecture. Extensive experiments on three benchmark datasets demonstrate significant improvements over state-of-the-art methods, yielding semantically consistent hash codes with superior retrieval performance.
This work addresses a key limitation in existing deep semantic hashing methods, where fixed-width and fixed-position semantic channels induce discontinuities in the loss function, thereby hindering optimization. To overcome this, the authors propose the Dynamic Semantic Channel Hashing (DSCH) loss, which dynamically adjusts both the position and scale of semantic channels to yield a smoother loss landscape and enhance hash code learning. Additionally, they introduce a tie-aware mean Average Precision (mAP) metric to more accurately evaluate retrieval performance under discrete Hamming distances. Evaluated across both cross-modal and single-modal settings on two benchmark datasets using two distinct architectures, DSCH significantly outperforms current state-of-the-art methods in 35 out of 40 tasks, achieving up to a 1.75 percentage point improvement in tie-aware mAP.
Approximate nearest neighbor (ANN) search remains a core challenge in large-scale cross-modal retrieval. This paper systematically surveys the early development of learning-based hashing methods, focusing on data-driven optimization of projection functions and quantization strategies to map high-dimensional features into compact binary codes enabling efficient similarity computation in Hamming space. Distinguishing itself from random hashing, the survey categorizes approaches into supervised, unsupervised, and semi-supervised paradigms, covering key directions including multi-bit encoding, adaptive thresholding, and cross-modal extensions. It distills their theoretical foundations and design principles, elucidating the fundamental trade-offs among accuracy, efficiency, and generalizability. Furthermore, it establishes a structured conceptual framework that clarifies the applicability boundaries and open challenges of early models. By doing so, the work provides both a theoretical reference and an evolutionary roadmap for future research on interpretable, robust, and multimodal hashing.
This work addresses the challenge in unsupervised fine-grained image hashing where neglecting collision resistance often causes semantically similar yet distinct samples to be mapped to identical hash codes. To this end, the paper proposes the CS3H framework, which explicitly models collision resistance for the first time in this task. Specifically, it introduces a normalized Hamming distance loss computed in a single forward pass to directly optimize similarity in Hamming space, and incorporates a collision-aware attention module to enhance the learning of rare yet discriminative local features. Extensive experiments demonstrate that CS3H significantly outperforms existing methods across multiple benchmark datasets, achieving higher retrieval accuracy and stronger collision resistance while maintaining low computational overhead.
This work addresses the high data acquisition cost of existing unsupervised cross-modal hashing methods, which typically rely on large-scale image-text pairs. To overcome this limitation, we propose Global-Neighborhood Aligned Hashing (GNAH), a novel approach that effectively transfers the semantic structure of vision-language foundation models into a compact binary Hamming space under limited paired data. GNAH integrates a prototype-anchored global alignment module with a contrastive stochastic neighborhood alignment module, jointly preserving global semantic consistency and local structural relationships to mitigate overfitting in sparse pairing scenarios. Extensive experiments demonstrate that GNAH significantly outperforms state-of-the-art unsupervised cross-modal retrieval methods under data-constrained settings, highlighting its practical utility.
This work addresses the challenge of recovering cross-model object correspondences when given only two partially overlapping embedding datasets produced by distinct black-box encoders. It identifies and leverages, for the first time, the cross-model isometric consistency in local geometric structures induced by contrastive learning encoders, and proposes a training-free, iterative geometric embedding hashing method. Guided by a small set of seed anchor points, the approach progressively expands correspondences across views by integrating distance-based hashing with Beta-Bernoulli Bayesian posterior aggregation. Experiments demonstrate that the method achieves high-precision and robust vector alignment across diverse encoder pairs, overlap ratios, and anchor conditions, successfully enabling applications such as vector database fusion and cross-model clustering.
This work addresses the computational inefficiency of long-context large language models during decoding, where self-attention incurs substantial overhead due to repeatedly processing an ever-growing key-value cache. To overcome this, the authors propose BinaryPC, a training-free, data-aware hashing-based sparse attention mechanism that introduces binary principal component analysis into attention computation for the first time. By constructing compact binary hash codes and hash functions that preserve intrinsic data structure, BinaryPC enables highly efficient approximate attention. Evaluated across multiple models and long-context benchmarks, BinaryPC achieves accuracy on par with full attention while delivering a 3.56× higher inference throughput than FlashAttention, significantly outperforming existing sparse and hashing-based baselines.
This work addresses the weak coupling between training objectives and discrete retrieval goals in existing cross-modal hashing methods, which typically learn semantics in continuous space and generate binary codes via sign functions. To overcome this limitation, the authors propose a unified framework based on spiking neural networks that formulates cross-modal hashing as a multi-timestep process involving spiking state evolution, directional spike interactions, and competitive spike readout. By replacing conventional continuous hashing heads with a positive-negative spike competition mechanism, the model directly optimizes image and text representations within the hash space, achieving strong alignment between training and retrieval. The proposed method attains competitive retrieval accuracy on three benchmark datasets while significantly reducing model parameters, computational cost, and energy consumption.