[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-137935-105":59,"doc-detail-137935-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","cdics-delving-into-fine-grained-attribute-for-in-context-segmentation-via-compositional-prompts-and-phased-decoupling","CDICS - Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling","","In-Context Learning (ICL) has shown strong performance for image segmentation by using reference images rather than text. However, selecting an ideal single exemplar for rare, complex concepts and expressing precise attribute-level segmentation remains difficult, since many approaches focus mainly on semantic or instance understanding. CDICS introduces compositional prompts and phased task decoupling to control In-Context Segmentation compositionally. It derives semantic, part, and color prompt components, performs coarse semantic localization first, then refines with compositional appearance prompts to match specified attributes. Experiments on reconstructed datasets and benchmarks validate superior performance, while extending ICL segmentation toward real-world fine-grained needs.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/cdics-delving-into-fine-grained-attribute-for-in-context-segmentation-via-compositional-prompts-and-phased-decoupling/137935/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/cdics-delving-into-fine-grained-attribute-for-in-context-segmentation-via-compositional-prompts-and-phased-decoupling/137935.png","ImageObject",300,407,{"name":92,"@type":93},"Melati","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-18","2026-08-23",true,{"@type":102,"interactionType":103,"userInteractionCount":39},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What limitation of existing ICL segmentation motivates CDICS?","Question",{"text":112,"@type":113},"Existing methods usually capture semantic or instance-level information well, but they struggle to flexibly control segmentation granularity and precisely represent attribute combinations from the input.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"How does CDICS use compositional prompts for segmentation control?",{"text":117,"@type":113},"CDICS builds prompts from semantic, part, and color components derived from reference prompts, combining them to dynamically define segmentation targets and enable compositional, attribute-controlled inference.",{"name":119,"@type":110,"acceptedAnswer":120},"What is the role of phased task decoupling in CDICS?",{"text":121,"@type":113},"CDICS uses a two-stage decoupled architecture: Stage 1 performs coarse-grained semantic localization, and Stage 2 refines results using appearance prompts based on color similarity and part features to produce fine-grained composed segmentation masks.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},137935,1787471087,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":39,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":52,"language":139,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":140,"faqs":141,"seo_title":142,"seo_description":67,"update_tm":129,"read_time":143},962085570644,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.  \nExcept for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.  \nCDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling  \nZhiyu Li 1,2 , Dianmo Sheng 1,2 , Qi Chu 1,2 , Shilong Chen 1,2 , Tao Gong 1,2†, Zhou Wei3 , Nenghai Yu 1,2  \n1University of Science and Technology of China 2Anhui Province Key Laboratory of Digital Security  \n3Lingyang Industrial Internet Co.,Ltd.  \n[linzeyin@mail.ustc.edu.cn](linzeyin@mail.ustc.edu.cn) , [tgong@ustc.edu.cn](tgong@ustc.edu.cn) , [ynh@ustc.edu.cn](ynh@ustc.edu.cn)  \nAbstract  \nIn-Context Learning (ICL) has shown great effectiveness in developing generalist image segmentation models. Its significant advantage over text-based descriptions is the ability to convey intricate visual appearance details through simple reference images. However, finding a perfectly matching single example for real-world rare and complex concepts is difficult. Moreover, existing methods are largely confined to semantic or instance-level understanding of the reference image, struggling to express more precise segmentation needs through the input. To address this, we propose CDICS, a novel framework that leverages Compositional prompts and phased task Decoupling to achieve compositional prompt-controlled In-Context Segmentation. Our method introduces compositional prompts derived from reference prompts, combining semantic, part and color images to dynamically define segmentation targets. To effectively fuse this control information, ensure synergy while suppressing interference, and mitigate feature coupling risks, our decoupled two-stage architecture firstly performs coarse-grained semantic localization, then refines the result using compositional appearance prompts to precisely match the specified attributes. This design extends traditional incontext segmentation, enabling it to support compositional prompts. Additionally, we reconstructed two datasets and their benchmarks to acquire data with part-color-specific attributes. Our method demonstrates superior performance on the compositional prompt-controlled in-context segmentation task. It also extends the capabilities of existing incontext segmentation, and makes an attempt toward realworld fine-grained segmentation.  \n1. Introduction  \nIn recent years, In-Context Learning (ICL) has emerged as a  \nprominent paradigm in image segmentation, enabling mod-† Corresponding author.  \nInput Output  \n\n| | Target Image Semantic&Instance-level\u003Cbr>|\n| --- | --- |\n\n\n| Ref Image semantic\u003Cbr>\u003Cbr>Ref Image Color\u003Cbr>| |\n| --- | --- |\n\nFigure 1 . Comparison of our two-stage CDICS framework with previous methods. The blue-highlighted area in Ref Image serves as a mask prompt. Previous methods perform standard semantic and instance-level segmentation based on “RefImage semantic”. Stage 1 of our method is similar to previous methods, while Stage 2 builds upon this foundation by applying additional appearance prompts (“RefImage Color” and “RefImage Part”) to achieve finegrained, compositional prompt-controlled segmentation.  \nels to segment novel targets by referencing just one or more exemplar images. Despite its efficacy, this paradigm faces a limitation: extant methods excel at understanding the semantic information of the target but struggle to flexibly adjust the granularity of segmentation based on the input.  \nIn many real-world scenarios, user’s needs are diverse. Users may sometimes require coarse-grained semantic segmentation, and at other times, fine-grained segmentation with precise appearance conditions such as “a person wearing a necklace of a specific style and color.” Furthermore, it is often impractical and burdensome for users to find exemplar images that perfectly match the desired attribute combination, highlighting the need for a more ","cbCaiawuLhe1gysw","https://ap.wps.com/l/cbCaiawuLhe1gysw","pdf",2488951,"English","# Introduction\n## Background: ICL-based image segmentation\n## Challenge: controlling granularity and attribute combinations\n## Proposed idea: compositional prompts and phased task decoupling\n## Method overview: two-stage encoder-decoder design","[{\"question\":\"What limitation of existing ICL segmentation motivates CDICS?\",\"answer\":\"Existing methods usually capture semantic or instance-level information well, but they struggle to flexibly control segmentation granularity and precisely represent attribute combinations from the input.\"},{\"question\":\"How does CDICS use compositional prompts for segmentation control?\",\"answer\":\"CDICS builds prompts from semantic, part, and color components derived from reference prompts, combining them to dynamically define segmentation targets and enable compositional, attribute-controlled inference.\"},{\"question\":\"What is the role of phased task decoupling in CDICS?\",\"answer\":\"CDICS uses a two-stage decoupled architecture: Stage 1 performs coarse-grained semantic localization, and Stage 2 refines results using appearance prompts based on color similarity and part features to produce fine-grained composed segmentation masks.\"}]","CDICS - Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decoupling | PDF",25]