[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84969-en":3,"doc-seo-84969-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84969,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Sparse Attention for Dense Open-Vocabulary Prediction in CLIP","Contrastive Language–Image Pre-training (CLIP) uses softmax-based self-attention whose strictly positive weights distribute probability mass to every token pair, including semantically irrelevant ones. This densification helps global context but introduces noise that obscures fine-grained, spatially localized cues needed for dense open-vocabulary prediction. This work replaces the row-wise final visual self-attention softmax with α-entmax at inference, acting as an implicit denoiser, and evaluates gains on dense segmentation and fine-grained retrieval. Benefits scale with how broadly baseline attention spreads off the target class.","arXiv :2607 .07135v2 [ cs .CV] 13 Jul 2026  \nSparse Attention for Dense Open-Vocabulary Prediction in CLIP  \nFatimah Zohra 1 Chen Zhao 1 ,2 Shuming Liu 1 Bernard Ghanem 1  \n1 King Abdullah University of Science and Technology (KAUST)  \n2 Harbin Institute of Technology, Shenzhen  \nAbstract. Contrastive Language–Image Pre-training (CLIP) relies on softmax-based self-attention, a strictly positive distribution that assigns probability mass to every pair of tokens—even semantically irrelevant ones. While these dense softmax weights are effective for gathering broad context during pre-training, they spread attention across many lowsalience tokens, producing noise that obscures the fine-grained, spatially localized cues required for dense, open-vocabulary prediction. We study an inference-time substitution of the row-wise softmax in the final visual self-attention layers with the α-entmax transform, applied across both the standard query–key attention and self-correlation variants. Because entmax applies a data-dependent threshold that maps low scores exactly to zero, it acts as an implicit denoiser, zeroing contextually irrelevant dependencies while redistributing mass onto the most relevant tokens. We evaluate on open-vocabulary tasks—dense semantic segmentation (Pascal VOC, Pascal Context, ADE20K) and fine-grained retrieval (FG-OVD)—and find the gain from attention sparsification is proportional to how much the baseline attention spreads off the target class.  \n1 Introduction  \nTransformer-based vision–language models such as CLIP [20] and its many largescale follow-ups, e.g., EVA-CLIP [7], SigLIP [30], and ALIGN [9], all use separate image and text encoders that are trained contrastively on hundreds of millions to billions of web-scraped image–text pairs, learning a shared embedding space for both modalities. This paradigm yields task-agnostic visual features that are aligned with language features, making CLIP the de facto front-end for zero-shot and open-vocabulary tasks.  \nHowever, CLIP’s global contrastive objective aligns a global representation of an image with a short caption, optimizing the visual encoder for image-level semantics rather than the spatially localized features that dense prediction tasks require. As a result, when CLIP’s patch tokens are used zero-shot for pixelor region-level tasks, their features localize poorly: patch-level analyses show that high activations frequently land on irrelevant background regions rather than the object of interest [21], and the attention maps of the final layers grow diffuse [3, 13] . The same image-level bias effects ROI-based fine-grained region  \n2 F. Zohra et al.  \nFig. 1: Attention Mass in Softmax vs. Entmax Distributions. Left: attention distribution for the query patch (green ×) under q-k attention. Entmax denoises the distribution, concentrating mass on the most relevant tokens –the horse’s body vs. the rider– while suppressing irrelevant background. Right: Sparsification helps in proportion to how much the baseline attention spreads off the target class. Where the mass is already very concentrated, entmax does not benefit. The x-axis shows the fraction of a foreground patch’s attention that lands on non-class (irrelevant) patches, measured against the downsampled Pascal VOC ground truth and the y-axis is the α-entmax (α= 1.2) minus softmax (α= 1.0) segmentation gain. α-entmax benefits when attention is spread across non-class tokens.  \nrecognition, where separating objects that differ by only a single attribute—color, material, or pattern—demands exactly the localized cues the global objective does not tease apart [14,29] .  \nA growing body of work improves CLIP’s dense predictions without retraining, by modifying the final attention block exclusively. The primary observations are that dropping the query–key mixing and keeping only the value path [5], substituting self-correlation (query–query, key–key, value–value) attention [3,13, 25], or reshaping the attention ne","cbCairvltaQmZqEx","https://ap.wps.com/l/cbCairvltaQmZqEx","pdf",24418977,1,21,"English","en",105,"# Introduction\n## Motivation: Dense softmax attention noise in CLIP\n## Related work: Dense prediction improvements by modifying attention\n## Core idea: α-entmax as inference-time sparsification\n## Evaluation setup and results overview","[{\"question\":\"Why does CLIP’s softmax attention hurt dense open-vocabulary prediction?\",\"answer\":\"Softmax attention assigns strictly positive weight to every token pair, spreading mass over low-salience, semantically irrelevant tokens. This diffuse attention can mask fine-grained, spatially localized cues required for dense prediction tasks.\"},{\"question\":\"What change does the paper make to CLIP at inference time?\",\"answer\":\"It substitutes the row-wise softmax in the final visual self-attention layers with the α-entmax transform. The replacement is applied to both standard query–key attention and self-correlation variants.\"},{\"question\":\"How does α-entmax produce the reported improvements?\",\"answer\":\"α-entmax applies a data-dependent threshold that maps low attention scores exactly to zero while preserving the probability simplex. This zeros contextually irrelevant dependencies and reallocates mass to the most relevant tokens, reducing denoising effects.\"}]",1784199827,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"sparse-attention-for-dense-open-vocabulary-prediction-in-clip","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/sparse-attention-for-dense-open-vocabulary-prediction-in-clip/84969/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does CLIP’s softmax attention hurt dense open-vocabulary prediction?","Question",{"text":75,"@type":76},"Softmax attention assigns strictly positive weight to every token pair, spreading mass over low-salience, semantically irrelevant tokens. This diffuse attention can mask fine-grained, spatially localized cues required for dense prediction tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What change does the paper make to CLIP at inference time?",{"text":80,"@type":76},"It substitutes the row-wise softmax in the final visual self-attention layers with the α-entmax transform. The replacement is applied to both standard query–key attention and self-correlation variants.",{"name":82,"@type":73,"acceptedAnswer":83},"How does α-entmax produce the reported improvements?",{"text":84,"@type":76},"α-entmax applies a data-dependent threshold that maps low attention scores exactly to zero while preserving the probability simplex. This zeros contextually irrelevant dependencies and reallocates mass to the most relevant tokens, reducing denoising effects.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]