[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86178-en":3,"doc-seo-86178-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86178,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","A Novel Method to Evaluate Models on Unreliable Noisy and Inconsistent Labels Adaptive Resolution Label Aggregation ARLA","Labels are central to training and evaluating deep learning segmentation models, yet real-world ground truth can be inconsistent, noisy, or ambiguous at class boundaries. This preprint introduces Adaptive Resolution Label Aggregation (ARLA), which adapts the spatial resolution of both labels and model predictions during inference before computing evaluation metrics. ARLA improves analysis of model behavior without retraining and is demonstrated on flood prediction, addressing inconsistent forest labels and cloud-cover label errors. Adjustable parameters align aggregated resolution with label precision and noise.","A NOVEL METHOD TO EVALUATE MODELS ON UNRELIABLE, NOISY AND INCONSISTENT LABELS:  \nADAPTIVE RESOLUTION LABEL AGGREGATION (ARLA)  \nPREPRINT  \n Natasha Randall  \nInstitute of Information Science Cologne University of Applied Sciences Claudiusstr. 1, Cologne, 50678, Germany [natasha.randall@th-koeln.de](natasha.randall@th-koeln.de)  \n Gernot Heisenberg  \nInstitute of Information Science Cologne University of Applied Sciences Claudiusstr. 1, Cologne, 50678, Germany [gernot.heisenberg@th-koeln.de](gernot.heisenberg@th-koeln.de)  \narXiv :2607 . 11214v1 [ cs .CV] 13 Jul 2026  \nABSTRACT  \nLabels are critical for both training and evaluating deep learning segmentation models, but are often inconsistent, noisy, or ambiguous at class boundaries. Many approaches have been developed to support training models on weak labels, but few to none currently exist to facilitate evaluating models on unreliable labels. We therefore introduce a method called ‘Adaptive Resolution Label Aggregation’, or ‘ARLA’, which dynamically adapts the resolution of both the label and the model prediction at inference time before the evaluation metrics are computed. We demonstrate how ARLAcan be used to better analyse model behaviour with a practical application to a real flood prediction model, where ARLA was able to overcome issues with inconsistent labelling of forested areas and errors in labels within regions of heavy cloud cover. Our work presents a new approach to evaluating segmentation models, with adjustable parameters to adapt the aggregated resolution to the precision of the label or the level of label noise. Fundamentally, ARLA exploits the information encapsulated by a label but minimises the label error, extracting from the noise a clearer signal of a model’s true performance.  \n1 Introduction  \nReliable ground truth labels are fundamental to the development of supervised machine learning or deep learning models, especially for segmentation or image-to-image prediction tasks [7] . However, labels can be very difficult or expensive to create [22], and thus are often generated using (partially) automated processes [21] . The gold standard for label creation is from manual human annotation, for example, drawing polygons by hand on remote sensing satellite data to distinguish different land cover classes [10] . Yet even seemingly gold standard labels frequently have issues; Northcutt et al. [19] estimated an average of at least 3.3% errors occurring across 10 popular computer vision benchmark datasets that included ImageNet, CIFAR and MNIST, and follow-up work by Wong et al. [25] demonstrated that high labeller disagreement led to uncertainty in the reliability and reproducibility of the labels in these datasets.  \nAlthough many successful methods have consequently been developed to support training models on weak or noisy labels [23], there is a significant research gap regarding methodologies for evaluating models on unreliable labels. This research gap is a problem, because Northcutt et al. [19] identified 2916 (6%) errors in the ImageNet validation subset, and showed how even small numbers of errors in test labels can lead to wrong model selection decisions when based on test accuracy benchmarks. Similarly, Lam and Stork [15] found that statistically, a classifier with a true error rate of 6% will report an error rate that is 15% higher, when only 1% of the testing labels are incorrect.  \nThe contribution of our work is a method to evaluate the performance of segmentation models on unreliable, noisy or inconsistent labels, called ‘Adaptive Resolution Label Aggregation’, or ‘ARLA’. In principle, ARLA works by adapting the resolution of both the label and the model prediction-in accordance with the expected precision of the labels-when calculating evaluation metrics at inference time, without needing to re-train the model. It thus extracts a much clearer  \nsignal of both the model’s performance and true error, without obscuring either the model’s streng","cbCaiuqsxRBKPUEv","https://ap.wps.com/l/cbCaiuqsxRBKPUEv","pdf",21383029,5,1,11,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does ARLA address in segmentation model evaluation?\",\"answer\":\"ARLA addresses the issue that ground-truth labels used for evaluation can be unreliable, noisy, inconsistent, or ambiguous at class boundaries, which distorts evaluation metrics and model selection decisions.\"},{\"question\":\"How does Adaptive Resolution Label Aggregation (ARLA) work?\",\"answer\":\"ARLA dynamically adapts the resolution of both the label and the model prediction at inference time, in accordance with the expected precision of the labels, and only then computes the evaluation metrics.\"},{\"question\":\"Does ARLA require retraining the model?\",\"answer\":\"No. ARLA is designed to be applied at inference time, extracting a clearer signal of model performance and true error without retraining.\"}]",1784209138,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-novel-method-to-evaluate-models-on-unreliable-noisy-and-inconsistent-labels-adaptive-resolution-label-aggregation-arla","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-novel-method-to-evaluate-models-on-unreliable-noisy-and-inconsistent-labels-adaptive-resolution-label-aggregation-arla/86178/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does ARLA address in segmentation model evaluation?","Question",{"text":76,"@type":77},"ARLA addresses the issue that ground-truth labels used for evaluation can be unreliable, noisy, inconsistent, or ambiguous at class boundaries, which distorts evaluation metrics and model selection decisions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Adaptive Resolution Label Aggregation (ARLA) work?",{"text":81,"@type":77},"ARLA dynamically adapts the resolution of both the label and the model prediction at inference time, in accordance with the expected precision of the labels, and only then computes the evaluation metrics.",{"name":83,"@type":74,"acceptedAnswer":84},"Does ARLA require retraining the model?",{"text":85,"@type":77},"No. ARLA is designed to be applied at inference time, extracting a clearer signal of model performance and true error without retraining.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]