[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82732-en":3,"doc-seo-82732-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82732,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Beyond Post-Quantization Native Hash Learning with a Dedicated Hash Token","Efficient large-scale image retrieval depends on compact binary representations that preserve semantic similarity while enabling fast Hamming-space search. Existing CNN- and ViT-based deep hashing pipelines mostly use post-quantization, learning continuous features first and producing binary codes only at the final hash projection or binarization stage. This late discretization creates a feature-to-code discrepancy. HashViT introduces a dedicated HASH token in the transformer, decomposed into a Hash Register for binary generation and a Semantic Workspace for continuous auxiliary semantics, refined layer by layer with a lightweight adapter and trained with unified supervision and quantization regularization.","Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token  \nXinze Liu, Ding Wang, Hengjie Zhu, and Dayan Wu  \narXiv :2607 .03328v 1 [ cs .CV] 3 Jul 2026  \nAbstract—Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing provides an appealing solution, but most existing CNN- and ViT-based methods still follow a post-quantization paradigm, where continuous visual features are first learned and binary codes are then produced by a terminal hash projection or binarization operation. This late code generation creates a feature-to-code discrepancy between the continuously optimized representation space and the discrete Hamming space used for retrieval. To address this limitation, we propose HashViT, a Vision Transformer framework for native hash token learning. Instead of treating hashing as a terminal readout, HashViT introduces a dedicated HASH token that serves as a persistent, hash-oriented retrieval state inside the transformer. The HASH token is structurally decomposed into a Hash Register for direct binary code generation and a Semantic Workspace for preserving auxiliary continuous semantics. To enable effective workspace-to-register interaction, we further design a lightweight Hash Refinement Adapter that progressively refines the Hash Register across transformer layers. As a result, binaryoriented representations are formed through token evolution within the backbone, rather than being abruptly induced by an output-level projection. HashViT is optimized with a unified objective that combines learnable semantic center supervision, classtoken similarity distillation, and quantization regularization, encouraging the HASH token to encode semantically structured and compact binary representations. Extensive experiments on three widely used benchmarks demonstrate that HashViT achieves state-of-the-art or highly competitive retrieval performance while preserving the efficiency of compact Hamming codes. The source code is available at [https://github.com/Xinze919/HashViT](https://github.com/Xinze919/HashViT).  \nIndex Terms—Deep image hashing, hash token, native hash learning.  \nI. INTRODUCTION  \nTHE rapid growth of multimedia data has made efficient  \nand accurate image retrieval a fundamental requirement in large-scale visual applications [1]–[3] . Among existing solutions, deep hashing methods [4]–[13] have attracted sustained attention because they represent images with compact binary codes, enabling low storage cost and efficient Hamming-distance-based search. Unlike traditional hashing methods based on hand-crafted features or data-independent projections [14]–[16], deep hashing jointly learns visual representations and hash functions in an end-to-end manner,  \nCorresponding author: Dayan Wu.  \nXinze Liu, Hengjie Zhu, and Dayan Wu are with the Institute of Information Engineering, CAS, Beijing 100190, China. Xinze Liu and Hengjie Zhu are also with the School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100190, China (e-mail: [liuxinze@iie.ac.cn](liuxinze@iie.ac.cn); [zhuhengjie@iie.ac.cn](zhuhengjie@iie.ac.cn); [wudayan@iie.ac.cn](wudayan@iie.ac.cn)).  \nDing Wang is with Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, Maryland 21218, USA (e-mail: [dwang141@jh.edu](dwang141@jh.edu)).  \nFig. 1. Motivation of HashViT. (a) Conventional post-quantization hashing first learns continuous backbone features and then maps them to polarized hash features through a terminal hash layer, causing an abrupt output-level transition from the continuous feature space to the Hamming space. (b) HashViT maintains a dedicated HASH token inside the transformer, whose Hash Register gradually evolves toward a polarized distribution across layers.  \nand has become a widely adopted paradigm for large-scale image retrieval. In particular, most existing methods focus on designing stro","cbCaisSCAJLgExk1","https://ap.wps.com/l/cbCaisSCAJLgExk1","pdf",3238318,3,1,13,"English","en",105,"# Introduction\n## Motivation and post-quantization limitation\n## Native hash token modeling\n## Early deep hashing background","[{\"question\":\"What limitation does post-quantization deep hashing introduce?\",\"answer\":\"It learns continuous visual features first, then generates binary codes only at a terminal hashing stage, causing a mismatch between the optimized continuous feature space and the discrete Hamming space used for retrieval.\"},{\"question\":\"How does HashViT generate hash codes differently from existing methods?\",\"answer\":\"HashViT models hashing as an internal, evolving retrieval state by introducing a dedicated HASH token in the transformer rather than treating hashing as a final readout.\"},{\"question\":\"What components and training signals does HashViT use for effective native hash learning?\",\"answer\":\"The HASH token is decomposed into a Hash Register for direct binary generation and a Semantic Workspace for auxiliary continuous semantics, with a Hash Refinement Adapter that progressively refines the register across layers and a unified objective combining semantic-center supervision, class-token similarity distillation, and quantization regularization.\"}]",1784182567,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-post-quantization-native-hash-learning-with-a-dedicated-hash-token","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-post-quantization-native-hash-learning-with-a-dedicated-hash-token/82732/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation does post-quantization deep hashing introduce?","Question",{"text":75,"@type":76},"It learns continuous visual features first, then generates binary codes only at a terminal hashing stage, causing a mismatch between the optimized continuous feature space and the discrete Hamming space used for retrieval.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does HashViT generate hash codes differently from existing methods?",{"text":80,"@type":76},"HashViT models hashing as an internal, evolving retrieval state by introducing a dedicated HASH token in the transformer rather than treating hashing as a final readout.",{"name":82,"@type":73,"acceptedAnswer":83},"What components and training signals does HashViT use for effective native hash learning?",{"text":84,"@type":76},"The HASH token is decomposed into a Hash Register for direct binary generation and a Semantic Workspace for auxiliary continuous semantics, with a Hash Refinement Adapter that progressively refines the register across layers and a unified objective combining semantic-center supervision, class-token similarity distillation, and quantization regularization.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]