[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-137724-105":59,"doc-detail-137724-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","approximate-size-targets-are-sufficient-for-accurate-semantic-segmentation-research-insights","Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation - Research Insights","","Extending image-level supervision to semantic segmentation using approximate relative object-size distributions enables off-the-shelf architectures to achieve strong accuracy without full pixel-precise masks. The method replaces binary class tags with categorical size targets and optimizes a zero-avoiding KL-divergence loss against the average prediction. Results on PASCAL VOC rely on new human annotations, while evaluations on COCO and medical data use synthetically corrupted size targets. Standard networks remain robust, and some classes outperform pixel-level supervision, which degrades under mask errors.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/approximate-size-targets-are-sufficient-for-accurate-semantic-segmentation-research-insights/137724/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/approximate-size-targets-are-sufficient-for-accurate-semantic-segmentation-research-insights/137724.png","ImageObject",300,407,{"name":92,"@type":93},"8796093062539","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-18","2026-08-22",true,{"@type":102,"interactionType":103,"userInteractionCount":34},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"How does the paper use image-level supervision for semantic segmentation?","Question",{"text":112,"@type":113},"It derives an image-level average prediction from per-pixel softmax outputs and trains the model using approximate object-size targets represented as categorical distributions.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What loss function is introduced for approximate size targets?",{"text":117,"@type":113},"The approach defines a size-target loss using zero-avoiding KL-divergence between the target size distribution v and the average prediction ̅S.",{"name":119,"@type":110,"acceptedAnswer":120},"How are approximate object sizes validated and tested across datasets?",{"text":121,"@type":113},"Validation on PASCAL VOC uses new human annotations of approximate object sizes, while experiments on COCO and medical data use synthetically corrupted size targets to study robustness.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},137724,1787439013,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":66,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":133,"file_id":134,"file_url":135,"file_type":136,"file_size":137,"view_count":34,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":52,"language":138,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":139,"faqs":140,"seo_title":141,"seo_description":67,"update_tm":129,"read_time":142},8796093062539,"Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation  \nXingye Fan University of Waterloo [x44fan@uwaterloo.ca](x44fan@uwaterloo.ca)  \nZhongwen (Rex) Zhang University of Waterloo [z889zhan@uwaterloo.ca](z889zhan@uwaterloo.ca)  \nYuri Boykov University of Waterloo [yboykov@uwaterloo.ca](yboykov@uwaterloo.ca)  \narXiv :2503 .06954v1 [ cs .CV] 10 Mar 2025  \nAbstract  \nThis paper demonstrates a surprising result for segmentation with image-level targets: extending binary class tags to approximate relative object-size distributions allows offthe-shelf architectures to solve the segmentation problem. A straightforward zero-avoiding KL-divergence loss for average predictions produces segmentation accuracy comparable to the standard pixel-precise supervision with full ground truth masks. In contrast, current results based on class tags typically require complex non-reproducible architectural modifications and specialized multi-stage training procedures. Our ideas are validated on PASCAL VOC using our new human annotations of approximate object sizes. We also show the results on COCO and medical data using synthetically corrupted size targets. All standard networks demonstrate robustness to the size targets’errors. For some classes, the validation accuracy is significantly better than the pixel-level supervision; the latter isnot robust to errors in the masks. Our work provides new ideas and insights on image-level supervision in segmentation and may encourage other simple general solutions to the problem.  \n1. Introduction  \nOur image-level supervision approach to semantic segmentation can be easily summarized using only a few standard notions. Soft-max predictions Sp = (S1p,..., SKp) generated by the segmentation model at any pixel p are categorical distributions over given K classes of objects (including background) present in the given set of training images. At any image, the average prediction is defined as  \n¯S = |~~ ~~1Ω| p Sp (1)  \nwhere Ω is a set of pixels p in the given image. The average prediction ¯S = (¯S1 ,   , ¯SK ) is also a categorical distribution over K classes. It can be seen as an image-level  \nFigure 1 . Semantic segmentation from image-level supervision: test results for training by (a) log-barriers (9) and (b) our approximate size targets (2) . Full GT-mask supervision results are in (c) .  \nprediction of the relative or normalized sizes (volume, area, or cardinality) of the image objects.  \nWe assume training images have approximate size targets represented by categorical distributions v = (vk) . For each such image, our size-target loss is defined as  \nLsize = KL (v∥¯S) = Xvk ln v~~¯~~Skk~~ ~~ (2)  \nk  \nwhere KL is Kullback–Leibler divergence. Figure 1(b) shows some test results for a generic segmentation network (WR38 backbone) trained on PASCAL VOC using only image-level supervision with approximate size targets (8%  \nmean relative errors) . Our total loss is very simple: it combines size-target loss (2) and a common CRF loss (3) .  \n1.1. Overview of weakly-supervised segmentation  \nBy weakly-supervised semantic segmentation we refer to all methods that do not use full pixel-precise ground truth masks for training. Such full supervision is overwhelmingly expensive for segmentation and is unrealistic for many practical purposes. There are many forms of weak supervision for semantic segmentation, e.g. using partial pixel-level ground truth defined by “seeds” [26, 32] or/and image-level supervision by class-tags [3, 21, 28] . It is also common to incorporate self-supervision based on various augmentation ideas and contrastive losses [10, 19, 34] .  \nLack of supervision also motivates unsupervised loss functions such as standard old-school regularization objectives for low-level segmentation or clustering. For example, many methods [7, 17, 34] use variants of K-means objective (squared errors) enforcing the compactness of each class representation. It is also very common to use CRF-based pai","cbCaio0OoScUCziZ","https://ap.wps.com/l/cbCaio0OoScUCziZ","pdf",6286818,"English","# Introduction\n## Overview of weakly-supervised segmentation","[{\"question\":\"How does the paper use image-level supervision for semantic segmentation?\",\"answer\":\"It derives an image-level average prediction from per-pixel softmax outputs and trains the model using approximate object-size targets represented as categorical distributions.\"},{\"question\":\"What loss function is introduced for approximate size targets?\",\"answer\":\"The approach defines a size-target loss using zero-avoiding KL-divergence between the target size distribution v and the average prediction ̅S.\"},{\"question\":\"How are approximate object sizes validated and tested across datasets?\",\"answer\":\"Validation on PASCAL VOC uses new human annotations of approximate object sizes, while experiments on COCO and medical data use synthetically corrupted size targets to study robustness.\"}]","Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation - Research Insights | PDF",25]