[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82925-en":3,"doc-seo-82925-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82925,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Repurposing CLIP to Localize at Pixel Level","Large-scale vision-language models such as CLIP deliver open-set localization at the image level, yet converting this ability to pixel-level dense prediction remains difficult because global feature biases distort attention. The paper presents CLIPix, which repurposes CLIP by tracing its classification process to derive object-specific attentive regions as pixel-level localization cues. A Noise-Resistant Correction strategy refines these cues for more accurate segmentation, while a Localization Embedding fuses localization and enriched detail for high-resolution results. Experiments on PASCAL and COCO validate state-of-the-art performance.","Repurposing CLIP to Localize at Pixel Level  \nJiaxiang Fang, Member, IEEE, Shiqiang Ma, Member, IEEE, Jing Wang, Siyu Chen, Fei Guo, Member, IEEE, and Shengfeng He, Senior Member, IEEE  \narXiv :2607 .05253v2 [ cs .CV] 7 Jul 2026  \nAbstract—Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level dense prediction poses challenges due to global feature biases. In this paper, we introduce CLIPix, a simple yet effective framework that “repurposes” CLIP to perform pixel-level localization. By tracing back CLIP’s classification process, CLIPix identifies object-specific attentive regions and repurposes them as pixellevel localization cues. To address noise introduced by global biases, we propose a Noise-Resistant Correction strategy, refining these cues for more precise segmentation. Additionally, we introduce a Localization Embedding strategy to integrate both localization and enriched detail information, enabling accurate, high-resolution segmentation. Our approach preserves CLIP’s generalization strength and unlocks its potential for segmenting arbitrary objects. Extensive experiments on the PASCAL and COCO datasets demonstrate that CLIPix achieves state-of-theart performance, underscoring its effectiveness. Our code is available at [github.com/aqingaqinghh/CLIPix](github.com/aqingaqinghh/CLIPix).  \nIndex Terms—Localize at pixel level, binary open-set semantic segmentation, noise-resistant correction, localization embedding.  \nI. INTRODUCTION  \nThe rise of large-scale datasets and enhanced computational power has fueled the development of expansive pre-trained models [1]–[3], renowned for their impressive generalization abilities [4] . These models have led to significant advancements in semantic segmentation, with particular attention to large visual-language models like CLIP [3] . While models such as the Segment Anything Model (SAM) encounter limitations, such as manual prompt inefficiencies and false positivesin automated prompting [5]–[7], CLIP offers a promising  \nThis work is supported by grants from the National Natural Science Foundation of China (Grants No. 62322215, 62532017 and 62402488), Natural Science Foundation of Hunan Province (Grants No. 2026JJ30018), the Guangdong Natural Science Funds for Distinguished Young Scholars (Grant 2023B1515020097), the National Research Foundation Singapore under the AI Singapore Programme (AISG Award No: AISG4-TC-2025-018-SGKR), and the Lee Kong Chian Fellowships. (Jiaxiang Fang and Shiqiang Ma contributed equally to this work.) (Corresponding authors: Fei Guo; Shengfeng He.)  \nJiaxiang Fang is with the School of Computer Science and Engineering, Central South University, Changsha 410083, China, and is with the Advanced Technology Center Beijing AI Laboratory, Chao-Yang District, Beijing 100027, China (e-mail: [254701041@csu.edu.cn](254701041@csu.edu.cn)) .  \nSiyu Chen and Fei Guo are with the School of Computer Science and Engineering, Central South University, Changsha 410083, China (e-mail: [csy619@csu.edu.cn](csy619@csu.edu.cn), [guofei@csu.edu.cn](guofei@csu.edu.cn)).  \nShiqiang Ma is with the Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China (e-mail: [sq.ma@siat.ac.cn](sq.ma@siat.ac.cn)).  \nJing Wang is with the Advanced Technology Center Beijing AI Laboratory, Chao-Yang District, Beijing 100027, China (e-mail: [jingd.wang@sony.com](jingd.wang@sony.com)).  \nShengfeng He is with the School of Computing and Information Systems, Singapore Management University, Singapore 188065 (e-mail: [shengfenghe@smu.edu.sg](shengfenghe@smu.edu.sg)).  \nalternative with higher efficiency and generalization potential [8] . By training on extensive paired image-caption datasets from the internet, CLIP effectively aligns visual and language spaces, enabling the model to locate objects based on object category names, thus addressing ","cbCaik6fxllY3PuV","https://ap.wps.com/l/cbCaik6fxllY3PuV","pdf",2822451,4,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does CLIPix address compared with using CLIP directly for segmentation?\",\"answer\":\"Direct adaptation of CLIP to pixel-level dense prediction is challenged by global feature biases that lead to inaccurate attention and false positives. CLIPix repurposes CLIP’s classification process to produce pixel-level localization cues and then corrects noise to improve segmentation accuracy.\"},{\"question\":\"How does CLIPix obtain pixel-level localization cues from CLIP?\",\"answer\":\"CLIPix traces back CLIP’s classification process to identify object-specific attentive regions. These regions are repurposed as pixel-level localization cues for dense prediction.\"},{\"question\":\"What strategies does CLIPix use to improve segmentation precision and resolution?\",\"answer\":\"It proposes a Noise-Resistant Correction strategy to refine cues affected by global biases. It also introduces a Localization Embedding strategy to integrate localization signals with enriched detail information for accurate, high-resolution segmentation.\"}]",1784184005,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"repurposing-clip-to-localize-at-pixel-level","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/repurposing-clip-to-localize-at-pixel-level/82925/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CLIPix address compared with using CLIP directly for segmentation?","Question",{"text":75,"@type":76},"Direct adaptation of CLIP to pixel-level dense prediction is challenged by global feature biases that lead to inaccurate attention and false positives. CLIPix repurposes CLIP’s classification process to produce pixel-level localization cues and then corrects noise to improve segmentation accuracy.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CLIPix obtain pixel-level localization cues from CLIP?",{"text":80,"@type":76},"CLIPix traces back CLIP’s classification process to identify object-specific attentive regions. These regions are repurposed as pixel-level localization cues for dense prediction.",{"name":82,"@type":73,"acceptedAnswer":83},"What strategies does CLIPix use to improve segmentation precision and resolution?",{"text":84,"@type":76},"It proposes a Noise-Resistant Correction strategy to refine cues affected by global biases. It also introduces a Localization Embedding strategy to integrate localization signals with enriched detail information for accurate, high-resolution segmentation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]