[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86110-en":3,"doc-seo-86110-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86110,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception","Open-vocabulary dense perception (OVDP) localizes objects unseen during training by using textual knowledge. Despite recent advances in CLIP-based methods, synonym-induced grounding inconsistency persists: semantically equivalent expressions generate different spatial attention patterns, hurting robustness and real-world performance. SynCLIP introduces a Synonym-Coherent Language-Image Pretraining framework with a Semantic-consistent Spatial Attention alignment (SSA) module and a Spatial Attention Refinement (SAR) module. Synonym-Enriched Visual Corpus (SEViC) supports training with multiple synonyms and definitions. Experiments show improved grounding consistency and state-of-the-art OVDP results.","SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust  \nOpen-Vocabulary Dense Perception  \nMingjie Xie 1 Guangjun He2 * Dongli Xu3 Youtian Lin4 Hongjue Li 1 Pengming Feng2 Jian Guan5 Yue Deng 1 ,6  \n1Beihang University 2 State Key Laboratory of Space Information System and Integrated Application  \n3Independent Researcher 4Nanjing University 5Harbin Engineering University 6Beijing Zhongguancun Academy  \narXiv :2607 . 11008v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nOpen-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynCLIP, a Synonym-Coherent Language-Image Pretraining framework that enhances synonym-robust grounding for OVDP. SynCLIP introduces a Semantic-consistent Spatial Attention alignment (SSA) module to enhance spatial attention consistency by minimizing discrepancies between attention maps of original and synonymous expressions. Furthermore, a Spatial Attention Refinement (SAR) module selectively strengthens the most semantically relevant spatial regions within aligned maps for more precise and stable grounding. To support synonym-coherent pretraining, we also construct a Synonym-Enriched Visual Corpus (SEViC), which augments each category with multiple synonymsand textual definitions. Extensive experiments on multiple benchmarks demonstrate that SynCLIP substantially improves grounding consistency under diverse linguistic variants and achieves state-of-the-art performance among CLIP-based OVDP methods. Code is available at [https:](https:)// [github.com/Justlovesmile/SynCLIP](github.com/Justlovesmile/SynCLIP).  \n1. Introduction  \nOpen-vocabulary dense perception (OVDP) aims to recognize and localize objects from novel categories that are not predefined in the training set by leveraging textual knowledge [46] . Unlike traditional dense perception tasks such  \nas object detection and segmentation [7, 10], which oper-*Corresponding author.  \nFigure 1 . Illustration of synonym-induced grounding inconsistency. (a) shows inconsistent dense perception across diverse synonymous expressions. (b) presents performance degradation caused by synonym-induced grounding inconsistency. Here,‘Novel’ and ‘Base’ denote the performance on unseen and seen categories, respectively.  \nate within a fixed label space and thus struggle to generalize to real-world environments with an open set of object categories [31, 35], OVDP employs textual expressions to flexibly represent labels. This language-driven supervision enables OVDP models to generalize beyond the closed-set categories when handling unseen categories. Consequently, OVDP has attracted growing attention for deployment in real-world applications such as robotics and autonomous driving [42, 45] .  \nTo enable open-vocabulary capability, recent studies [4, 38, 43] commonly leverage vision-language models (VLMs) pretrained on large-scale image–text pairs, such as CLIP [17, 24, 26] . These methods transfer the global vision–language alignment learned by VLMs to region-level representations, allowing visual regions to be associated with category-related textual expressions and thereby facilitating the recognition and localization of unseen categories. For example, a series of approaches [30, 33, 34] enhance  \nFigure 2 . Comparison of spatial attention distributions generated by different models for synonymous expressions, e.g., synonymsand definition. Our method yields more consistent and semantically aligned attention, indicating synonym-robust grounding.  \nregion-level vision–language alignment by employing techniques such as pseudo-region self-dis","cbCait7Nr6mow578","https://ap.wps.com/l/cbCait7Nr6mow578","pdf",30059924,3,1,15,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does SynCLIP address in open-vocabulary dense perception?\",\"answer\":\"SynCLIP targets synonym-induced grounding inconsistency, where semantically equivalent expressions produce inconsistent spatial attention and degraded localization performance.\"},{\"question\":\"How does SynCLIP improve synonym-robust grounding?\",\"answer\":\"It uses a Semantic-consistent Spatial Attention alignment (SSA) module to align attention maps across original and synonymous expressions, and a Spatial Attention Refinement (SAR) module to strengthen the most semantically relevant spatial regions.\"},{\"question\":\"What is the role of SEViC in training SynCLIP?\",\"answer\":\"SEViC constructs a synonym-enriched visual corpus by augmenting each category with multiple synonyms and textual definitions, enabling synonym-coherent pretraining.\"}]",1784208585,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"synclip-synonym-coherent-language-image-pretraining-for-robust-open-vocabulary-dense-perception","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/synclip-synonym-coherent-language-image-pretraining-for-robust-open-vocabulary-dense-perception/86110/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SynCLIP address in open-vocabulary dense perception?","Question",{"text":75,"@type":76},"SynCLIP targets synonym-induced grounding inconsistency, where semantically equivalent expressions produce inconsistent spatial attention and degraded localization performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SynCLIP improve synonym-robust grounding?",{"text":80,"@type":76},"It uses a Semantic-consistent Spatial Attention alignment (SSA) module to align attention maps across original and synonymous expressions, and a Spatial Attention Refinement (SAR) module to strengthen the most semantically relevant spatial regions.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of SEViC in training SynCLIP?",{"text":84,"@type":76},"SEViC constructs a synonym-enriched visual corpus by augmenting each category with multiple synonyms and textual definitions, enabling synonym-coherent pretraining.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]