[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85083-en":3,"doc-seo-85083-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85083,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation","Referring Image Segmentation (RIS) segments image regions guided by natural-language descriptions, enabling controllable, fine-grained visual understanding. Adapting RIS to endoscopy is difficult due to scarce high-quality annotations and intricate domain-specific image–text relationships, and existing vision–language models often miss subtle textual cues, limiting accuracy and generalization. ReferEndoscopy provides a large-scale endoscopy benchmark for RIS, and AR-ERIS uses attribute retrieval with open-vocabulary compositional referring segmentation. Pretrained on ReferEndoscopy, it achieves state-of-the-art results and strong transfer across simulated and real endoscopic data.","ATTRIBUTE RETRIEVING FOR OPEN-VOCABULARY ENDOSCOPIC COMPOSITIONAL REFERRING SEGMENTATION  \nShun Liu  \nVirginia Commonwealth University  \n[lius24@vcu.edu](lius24@vcu.edu)  \nTianyu Luan  \nUniversity at Buffalo [tianyulu@buffalo.edu](tianyulu@buffalo.edu)  \nNan Xi  \nVirginia Commonwealth University  \n[xin@vcu.edu](xin@vcu.edu)  \nYang Liu  \nKing’s College London [yang.9.liu@kcl.ac.uk](yang.9.liu@kcl.ac.uk)  \nXuan Gong  \nHarvard Medical School [xuan_gong@hms.edu](xuan_gong@hms.edu)  \nDavid Doermann  \nUniversity at Buffalo  \n[doermann@buffalo.edu](doermann@buffalo.edu)  \narXiv :2607 .08397v 1 [ cs .CV] 9 Jul 2026  \nABSTRACT  \nReferring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique challenges, including scarce high-quality annotations and complex, domainspecific image–text relationships. Although recent vision–language models demonstrate strong cross-domain alignment, they often fail to capture fine-grained textual cues in endoscopic settings, resulting in suboptimal performance and limited generalization. To address these challenges, we introduce ReferEndoscopy, a large-scale benchmark for RIS in the endoscopy field. Building on this dataset, we propose the Attribute Retrieval-based Endoscopic-RIS (AR-ERIS) framework for open-vocabulary endoscopic compositional referring segmentation. AR-ERIS leverages attribute retrieval for open-vocabulary endoscopic compositional referring segmentation and is pretrained on the curated ReferEndoscopy dataset, achieving state-of-the-art performance with strong generalization across both simulated and real-world endoscopic data. The dataset and code will be publicly released upon completion of the review process.  \n1 Introduction  \nAnalyzing endoscopic images is vital in minimally invasive surgeries and diagnostics by providing real-time, highresolution visualization of internal structures and organs. Precise identification and segmentation of anatomical regions and surgical instruments are essential to support clinical decision-making and enhance procedural outcomes. While traditional segmentation methods [1, 2, 3] have achieved notable success in general biomedical image segmentation, endoscopic image segmentation poses unique challenges due to limited detailed annotations, complex textures, occlusions, and dynamic interactions between instruments and tissues.  \nReferring image segmentation (RIS) [4], which segments object regions based on textual descriptions, has gained recent attention in natural image segmentation. Extending RIS to the medical domain, especially for endoscopic images [5, 6, 7, 8], offers significant potential but remains largely unexplored. As shown in Figure 1, in intelligent surgical systems, RIS can be applied to segment specific objects based on textual prompts provided by a user, enhancing the system’s responsiveness to real-time instructions. Unlike static objects in natural images, surgical instruments and tissues in endoscopic videos move and interact frequently, often under challenging conditions such as low lighting, smoke, occlusion, truncation, and blurred boundaries. Efficiently capturing boundary and shape information for open-set categories is thus a critical research question. Current endoscopic image segmentation methods [7, 8] rely primarily on image-only data within a limited domain, lacking the textual guidance needed for better generalization. This gap highlights the urgent need for a comprehensive text-grounded benchmark and a robust referring segmentation approach. An effective RIS model should accurately segment anatomical structures and instruments in real time based on textual prompts, enhancing the precision and situational awareness of surgical support systems.  \nAttribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation  \nIntelligent surgical System  \n“The ca","cbCaisVMPYT94lMR","https://ap.wps.com/l/cbCaisVMPYT94lMR","pdf",3029460,1,14,"English","en",105,"# Abstract\n# Introduction\n# Attribute Retrieval for Open-Vocabulary Endoscopic Compositional Referring Segmentation\n## Application of RIS in Intelligent Surgical Systems\n## Proposed AR-ERIS Framework","[{\"question\":\"What problem does open-vocabulary endoscopic compositional referring segmentation address?\",\"answer\":\"It targets RIS in endoscopic images, where a model must segment regions specified by detailed natural-language prompts, supporting fine-grained and controllable visual understanding in surgical contexts.\"},{\"question\":\"Why is endoscopic RIS harder than RIS in natural images?\",\"answer\":\"Endoscopy has limited detailed annotations and complex relationships between images and text, plus challenging conditions such as low lighting, occlusions, truncation, and blurred boundaries that complicate boundary and shape capture.\"},{\"question\":\"How do ReferEndoscopy and AR-ERIS improve performance and generalization?\",\"answer\":\"ReferEndoscopy provides a large-scale text-grounded endoscopy benchmark, while AR-ERIS leverages attribute retrieval for open-vocabulary compositional referring segmentation and is pretrained on this curated dataset, yielding state-of-the-art results with strong transfer across simulated and real data.\"}]",1784200938,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"attribute-retrieving-for-open-vocabulary-endoscopic-compositional-referring-segmentation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/attribute-retrieving-for-open-vocabulary-endoscopic-compositional-referring-segmentation/85083/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does open-vocabulary endoscopic compositional referring segmentation address?","Question",{"text":75,"@type":76},"It targets RIS in endoscopic images, where a model must segment regions specified by detailed natural-language prompts, supporting fine-grained and controllable visual understanding in surgical contexts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is endoscopic RIS harder than RIS in natural images?",{"text":80,"@type":76},"Endoscopy has limited detailed annotations and complex relationships between images and text, plus challenging conditions such as low lighting, occlusions, truncation, and blurred boundaries that complicate boundary and shape capture.",{"name":82,"@type":73,"acceptedAnswer":83},"How do ReferEndoscopy and AR-ERIS improve performance and generalization?",{"text":84,"@type":76},"ReferEndoscopy provides a large-scale text-grounded endoscopy benchmark, while AR-ERIS leverages attribute retrieval for open-vocabulary compositional referring segmentation and is pretrained on this curated dataset, yielding state-of-the-art results with strong transfer across simulated and real data.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]