Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token

📅 2026-07-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing deep hashing methods are limited by the representational discrepancy between continuous features and discrete Hamming codes. This work proposes HashViT, the first framework to natively embed hash learning within a Vision Transformer by introducing a dedicated HASH token that progressively evolves binary codes layer-by-layer inside the network, thereby circumventing end-of-pipeline quantization. The HASH token comprises a Hash Register and a Semantic Workspace, complemented by a lightweight Hash Refinement Adapter for fine-grained optimization. Through a joint training strategy integrating learnable semantic centroid supervision, class-token similarity distillation, and quantization regularization, HashViT achieves state-of-the-art or highly competitive performance on three mainstream image retrieval benchmarks while preserving the efficiency advantages of compact Hamming codes for fast retrieval.
📝 Abstract
Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing is appealing, but most existing CNN- and ViT-based methods still follow a post-quantization paradigm, where continuous visual features are first learned and binary codes are then produced by a terminal hash projection or binarization operation. This late code generation creates a feature-to-code discrepancy between the continuously optimized representation space and the discrete Hamming space used for retrieval. To address this limitation, we propose HashViT, a Vision Transformer framework for native hash token learning. Instead of treating hashing as a terminal readout, HashViT introduces a dedicated HASH token that serves as a persistent, hash-oriented retrieval state inside the transformer. The HASH token is structurally decomposed into a Hash Register for direct binary code generation and a Semantic Workspace for preserving auxiliary continuous semantics. To enable effective workspace-to-register interaction, we further design a lightweight Hash Refinement Adapter that progressively refines the Hash Register across transformer layers. As a result, binary-oriented representations are formed through token evolution within the backbone, rather than being abruptly induced by an output-level projection. HashViT is optimized with a unified objective that combines learnable semantic center supervision, class-token similarity distillation, and quantization regularization, encouraging the HASH token to encode semantically structured and compact binary representations. Extensive experiments on three widely used benchmarks demonstrate that HashViT achieves state-of-the-art or highly competitive retrieval performance while preserving the efficiency of compact Hamming codes. Code is available at https://github.com/Xinze919/HashViT.
Problem

Research questions and friction points this paper is trying to address.

deep hashing
post-quantization
binary code
Hamming space
feature-to-code discrepancy
Innovation

Methods, ideas, or system contributions that make the work stand out.

native hash learning
HASH token
Vision Transformer
binary code generation
post-quantization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xinze Liu
Institute of Information Engineering, CAS, Beijing 100190, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100190, China
D
Ding Wang
Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, Maryland 21218, USA
H
Hengjie Zhu
Institute of Information Engineering, CAS, Beijing 100190, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100190, China
D
Dayan Wu
Institute of Information Engineering, CAS, Beijing 100190, China