[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81703-en":3,"doc-seo-81703-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81703,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps","Advanced foundation models such as ModernBERT deliver stronger dense retrieval than earlier architectures, yet they underperform established BERT-base baselines in learned sparse retrieval (LSR). The paper attributes this gap to a Vocabulary Gap: modern tokenizers use raw, case-sensitive vocabularies that map single semantics to redundant surface forms, wasting capacity on morphological noise and reducing lexical matching quality. It introduces a theoretical framework and Vocabulary Transfer (VT) to normalize vocabularies while preserving semantic integrity and aligning manifolds with sparsity constraints.","Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps  \narXiv :2607 .00004v1 [ cs .IR] 20 Apr 2026  \nZhichao Geng  \nAmazon Web Service  \nShanghai, China [zhichaog@amazon.com](zhichaog@amazon.com)  \nAbstract  \nWhile advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the aging BERT-base baseline in learned sparse retrieval (LSR) . We identify the root cause as the Vocabulary Gap: modern tokenizers utilize raw, case-sensitive vocabularies designed for lossless reconstruction, which map single semantic units to redundant surface forms, wasting model capacity on morphological noise and hindering lexical matching. We formalize this intuition through a theoretical framework, demonstrating that appropriate vocabulary coarse-graining can tighten the generalization bounds by reducing complexity of the hypothesis class, provided that semantic integrity is preserved. To resolve this, we propose Vocabulary Transfer (VT), a model-agnostic framework that migrates advanced encoders to sparse-friendly, normalized vocabularies with minimal computational cost. VT utilizes a novel Semantic Initialization via spatial topology to preserve geometric structure and an Activation Potential Calibration (APC) mechanism to align pre-trained manifolds with sparsity constraints, preventing the dead neuron and dense collapse observed in standard fine-tuning. Empirically, VT is universally effective: it enables ModernBERT to achieve state-ofthe-art performance on the BEIR benchmark (52.4 nDCG, a +4.7 improvement), resuscitates failing models like RoBERTa-large, and generalizes seamlessly to inference-free architectures and specialized domains. These results confirm that the performance lag isnot an architectural deficiency but a solvable vocabulary mismatch. We’ve released our code and models.1  \nCCS Concepts  \n• Information systems → Retrieval models and ranking.  \nKeywords  \nSPLADE, learned sparse representations, passage retrieval  \nACM Reference Format:  \nZhichao Geng and Yang Yang. 2026. Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’26), July 20–24, 2026, Melbourne, VIC, Australia. ACM, New York, NY, USA, 12 pages. [https://doi.org/10.1145/](https://doi.org/10.1145/)[ ](https://doi.org/10.1145/)3805712.3809724  \n1[https://anonymous.4open.science/r/vocab-transfer/. All details](https://anonymous.4open.science/r/vocab-transfer/. All details) included.  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License. SIGIR ’26, Melbourne, VIC, Australia  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2599-9/2026/07  \n[https://doi.org/10.1145/3805712.3809724](https://doi.org/10.1145/3805712.3809724)  \nYang Yang  \nAmazon Web Service Shanghai, China [yych@amazon.com](yych@amazon.com)  \nPerformance Gap with BERT-uncased (nDCG@10)  \n+2  \n+0  \n-2  \n-4  \n-10  \n-20  \n-40  \nDense Retrieval Sparse Retrieval Sparse Retrieval  \n(Input Lowercased)  \nFigure 1: The Vocabulary Gap anomaly. While advanced encoders like ModernBERT significantly outperform BERT in dense retrieval, they lag behind in sparse retrieval under standard fine-tuning.  \n1 Introduction  \nThe landscape of neural information retrieval has bifurcated into two dominant paradigms: dense retrieval, which encodes queries and documents into continuous low-dimensional embeddings [21, 56], and learned sparse retrieval (LSR), which projects text into highdimensional, weighted lexical vectors [15, 33] . While dense retrievers excel at capturing semantic nuances, sparse retrievers—exemplified by models like SPLADE [15]—retain the interpretability and efficiency of inverted indices while mitigating the lexical mismatch problem of traditional BM25 [34, 47] .  \nIn","cbCaib0Uqy3x4TJ9","https://ap.wps.com/l/cbCaib0Uqy3x4TJ9","pdf",888968,2,1,12,"English","en",105,"# Introduction\n## Dense vs. learned sparse retrieval\n## Performance anomaly under sparse retrieval\n## Root cause: Vocabulary Gap\n## Vocabulary Transfer (VT) approach","[{\"question\":\"Why do advanced encoders lag in learned sparse retrieval even when they excel in dense retrieval?\",\"answer\":\"The paper links the degradation to the Vocabulary Gap: modern tokenization yields raw, redundant surface forms that hinder lexical matching and waste model capacity under sparse objectives.\"},{\"question\":\"What is the Vocabulary Gap, according to the document?\",\"answer\":\"It is the mismatch between modern tokenizers’ non-normalized, case-sensitive vocabularies (aimed at lossless reconstruction) and the lexical matching needs of learned sparse retrieval.\"},{\"question\":\"How does the proposed Vocabulary Transfer (VT) framework address the issue?\",\"answer\":\"VT migrates advanced encoders to sparse-friendly normalized vocabularies with minimal compute cost, using Semantic Initialization via spatial topology and Activation Potential Calibration to prevent dead neurons and dense collapse.\"}]",1784175506,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"why-advanced-encoders-lag-on-sparse-retrieval-the-answer-and-an-approach-to-bridging-vocabulary-gaps","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/why-advanced-encoders-lag-on-sparse-retrieval-the-answer-and-an-approach-to-bridging-vocabulary-gaps/81703/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do advanced encoders lag in learned sparse retrieval even when they excel in dense retrieval?","Question",{"text":75,"@type":76},"The paper links the degradation to the Vocabulary Gap: modern tokenization yields raw, redundant surface forms that hinder lexical matching and waste model capacity under sparse objectives.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Vocabulary Gap, according to the document?",{"text":80,"@type":76},"It is the mismatch between modern tokenizers’ non-normalized, case-sensitive vocabularies (aimed at lossless reconstruction) and the lexical matching needs of learned sparse retrieval.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed Vocabulary Transfer (VT) framework address the issue?",{"text":84,"@type":76},"VT migrates advanced encoders to sparse-friendly normalized vocabularies with minimal compute cost, using Semantic Initialization via spatial topology and Activation Potential Calibration to prevent dead neurons and dense collapse.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]