[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83344-en":3,"doc-seo-83344-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83344,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks","Leaderboard comparisons drive claims of progress in colonoscopy polyp segmentation, yet verification remains difficult. A systematic audit of 27 papers from 2015 to 2026 identifies three structural evaluation issues. Most papers omit Hausdorff distance, even though it directly reflects boundary accuracy important for detecting flat or small polyps. Multiple incompatible train/test split protocols coexist across the same datasets, making Dice scores incomparable. Performance claims also lack statistical significance tests. Controlled re-evaluation shows large boundary and recall failures hidden by reported metrics, metric-dependent winners, and unstable near-tied rankings across random splits. A five-point Polyp Segmentation Reporting Checklist (PSRC) is proposed to improve domain reporting.","arXiv :2607 .08203v 1 [ cs .CV] 9 Jul 2026  \nMetrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks  \nAisha Urooj 1 , Zain Ul Abdien2 , and Neelu Madan3  \n1 Lunit, South Korea  \n2 RPTU Kaiserslautern, Germany  \n3 Aalborg University & Pioneer Center, Denmark  \nAbstract. Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this apparent progress is difficult to verify: a systematic audit of 27 papers published between 2015 and 2026 reveals three structural problems in how the community evaluates models. First, 25 of 27 papers omit the Hausdorff distance. Hausdorff distance is a boundary-accuracy metric with direct clinical relevance for detecting flat or small polyps, and is a standard in radiotherapy segmentation.  \nSecond, at least five incompatible train/test split protocols co-exist across papers reporting results on the same two datasets (Kvasir-SEG and CVCClinicDB), making published Dice scores non-comparable even when they appear in the same leaderboard column. Third, 26 of 27 papers make performance claims without any statistical significance test. Strikingly, four papers published after the Metrics Reloaded framework [14](Maier-Heinet al. , Nature Methods 2024) perpetuate these same problems, suggesting that general-purpose metric guidance has not yet reached the colonoscopy sub-community. To show these problems are not merely cosmetic, were-evaluate five representative models under three controlled protocols with a single uniform scorer, and find that the reported metric conceals large boundary and recall failures, that the “best”model changes with the metric, and that near-tied rankings reverse across random splits. We propose a five-point Polyp Segmentation Reporting Checklist (PSRC) as a lightweight, domain-adapted corrective.  \nKeywords: evaluation metrics · polyp segmentation · reproducibility · benchmark audit · clinical metrics  \n1 Introduction  \nColorectal cancer is the third most prevalent cancer worldwide, and colonoscopy remains the primary screening modality [6] . However, manual examination is highly prone to human error, with up to 20-25% of precancerous polyps missed during routine procedures due to visual fatigue or subtle lesion morphology [20,31] . Because early detection and removal of these adenomas significantly lowers mortality rates, automated polyp segmentation can assist endoscopists by highlighting  \n2 Aisha Urooj, Zain Ul Abdien, and Neelu Madan  \nFig. 1. The task. Binary polyp segmentation maps a colonoscopy frame to a per-pixel polyp mask via an encoder–decoder network with deep supervision. The two clinically decisive failure modes are exactly the ones region-overlap metrics under-report: a missed polyp is a missed cancer (recall), and an inaccurate boundary is a wrong resection margin (HD95/NSD) .  \nlesions in real time. Consequently, the past five years have seen a rapid succession of deep learning models that each claim state-of-the-art (SOTA) performance on standard benchmarks.  \nYet a closer inspection of how SOTA is evaluated reveals a troubling pattern. The field converged around a single evaluation template introduced by PraNet [6] in 2020, where six metrics are reported across five datasets using a fixed train/test split. The metrics used in this template are: Mean Dice, Mean IoU, Weighted F-Measures, Structure Measure, Mean Absolute Error, and Enhanced-alignment Measure. The datasets used in the template are: Kvasir-SEG [8], CVC-ClinicDB/CVC-612 [2], CVC-ColonDB [25], ETIS-LaribPolypDB [23], EndoScene [28] This template is then subsequently copied, sometimes verbatim, by more than a dozen subsequent papers (e.g., ColonFormer [5], Polyp-PVT [4], SA-Net [29]) without questioning whether those six metrics adequately capture what matters clinically.  \nFrom a clinical standpoint, pixel-level overlap metrics like Mean Dice and Mean IoU remain in","cbCaikuGBrvXKaFg","https://ap.wps.com/l/cbCaikuGBrvXKaFg","pdf",5475779,2,1,19,"English","en",105,"# Introduction\n## Core problem: evaluation template vs clinical goals\n## Metric choice and boundary accuracy\n## Inconsistencies in data partitioning\n## Lack of statistical significance\n## Contributions and proposed PSRC","[{\"question\":\"What three structural evaluation problems are identified in the audit of polyp segmentation papers?\",\"answer\":\"The audit finds that most papers omit Hausdorff distance, that incompatible train/test split protocols are used for the same datasets, and that performance claims are often made without statistical significance testing.\"},{\"question\":\"Why is Hausdorff distance considered important for clinical relevance in polyp segmentation?\",\"answer\":\"Hausdorff distance provides boundary-accuracy information, which directly matters for detecting flat or small polyps and for clinically relevant resection-margin errors.\"},{\"question\":\"How can metric choice and split protocol affect the validity of benchmark comparisons?\",\"answer\":\"Incompatible split protocols make published Dice scores non-comparable, and controlled re-evaluation shows that near-tied rankings can reverse across random splits and that the “best” model can change depending on the metric.\"}]",1784186910,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"metrics-or-mirage-an-audit-of-evaluation-inconsistencies-in-colonoscopy-polyp-segmentation-benchmarks","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/metrics-or-mirage-an-audit-of-evaluation-inconsistencies-in-colonoscopy-polyp-segmentation-benchmarks/83344/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What three structural evaluation problems are identified in the audit of polyp segmentation papers?","Question",{"text":75,"@type":76},"The audit finds that most papers omit Hausdorff distance, that incompatible train/test split protocols are used for the same datasets, and that performance claims are often made without statistical significance testing.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is Hausdorff distance considered important for clinical relevance in polyp segmentation?",{"text":80,"@type":76},"Hausdorff distance provides boundary-accuracy information, which directly matters for detecting flat or small polyps and for clinically relevant resection-margin errors.",{"name":82,"@type":73,"acceptedAnswer":83},"How can metric choice and split protocol affect the validity of benchmark comparisons?",{"text":84,"@type":76},"Incompatible split protocols make published Dice scores non-comparable, and controlled re-evaluation shows that near-tied rankings can reverse across random splits and that the “best” model can change depending on the metric.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]