[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121581-en":3,"doc-seo-121581-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121581,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Mislabeled examples detection viewed as probing machine learning models - concepts, survey and extensive benchmark","Mislabeled examples are widespread in real-world machine learning datasets, motivating techniques for automatic detection. The work reframes most mislabeled detection methods as probing mechanisms applied to trained machine learning models, grounded in a small set of core principles. It formalizes a modular framework with four building blocks and provides a Python library implementation. Experiments target classifier-agnostic ideas and evaluate deep-learning-to-tabular adaptations. A benchmark covers NCAR and NNAR labeling noise across multiple tasks, offering new insights and limitations under imperfect labeling rules.","arXiv :2410 . 15772v1 [ cs .LG] 21 Oct 2024  \nMislabeled examples detection viewed as probing machine learning models: concepts, survey and extensive benchmark  \nThomas George∗ [thomas.george@orange. com](thomas.george@orange. com)  \nOrange Innovation  \nPierre Nodet∗ [pierre. nodet@orange. com](pierre. nodet@orange. com)  \nOrange Innovation  \nAlexis Bondu [alexis. bondu@orange. com](alexis. bondu@orange. com)  \nOrange Innovation  \nVincent Lemaire [vincent.lemaire@orange. com](vincent.lemaire@orange. com)  \nOrange Innovation  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= 3YlOr7BHkx](https: // openreview. net/ forum? id= 3YlOr7BHkx)  \nAbstract  \nMislabeled examples are ubiquitous in real-world machine learning datasets, advocating the development of techniques for automatic detection. We show that most mislabeled detection methods can be viewed as probing trained machine learning models using a few core principles. We formalize a modular framework that encompasses these methods, parameterized by only 4 building blocks, as well as a Python library that demonstrates that these principles can actually be implemented. The focus is on classifier-agnostic concepts, with an emphasis on adapting methods developed for deep learning models to non-deep classifiers for tabular data. We benchmark existing methods on (artificial) Completely At Random (NCAR) as well as (realistic) Not At Random (NNAR) labeling noise from a variety of tasks with imperfect labeling rules. This benchmark provides new insights as well as limitations of existing methods in this setup.  \n1 Introduction  \nIn supervised machine learning, the performance of learned algorithms crucially depends on the quality of the dataset of examples used during training: how many examples do we have access to, are these examples representative of the actual distribution on the feature space, and were the training examples correctly labeled? We focus on the latter subject. Indeed, many actual use cases include some amount of labeling errors. For example, this is typically the case in tasks that involve human supervision since labeling large datasets requires a pool of annotators that possess a mix of expert knowledge (which is costly), and willingness to perform repetitive tasks (which is dull) . This is also known to be the case for widely used benchmark datasets such as CIFAR10/100 or MNIST (Northcutt et al., 2021b) . Therefore, cleansing datasets offers the promise of better performance, but at the cost of additional efforts. Since the early days of machine learning, it has been widely believed that this could be achieved through automated methods, eliminating the need for further human intervention. This has led to many methods for automatic detection of mislabeled examples using classical machine learning methods (Guan & Yuan, 2013) . With the success of deep learning methods in applications ranging from image recognition to language models, new mislabeled detection methods have also been proposed that exploit its specific training dynamics.  \n*equal contribution  \nThe aim of this paper is to offer a new perspective on existing mislabeled detection methods, as well as practical recommendations in actual use cases in the presence of labeling noise. Rather than learning a model that captures the structure of the labeling noise, our approach is to blindly evaluate existing methods on real-world datasets, with no prior knowledge. We survey mislabeled detection methods regardless of whether they were designed to work with deep learning models or other classical machine learning algorithms, and we highlight a few common principles. We also focus on tabular and text data, a type of data that is prevalent in the industry (e.g. in logs, in customer databases, etc) but that has recently received less attention than datasets that are more amenable to deep learning methods such as images, sound, or language.  \nThis paper is organized as follows: In section 2, we suggest a d","cbCaid8UwPNuoLkv","https://ap.wps.com/l/cbCaid8UwPNuoLkv","pdf",3115844,1,43,"English","en",105,"# Abstract\n# 1 Introduction\n# List of contributions","[{\"question\":\"What is the core problem addressed in this paper?\",\"answer\":\"The paper studies how to detect mislabeled examples in supervised machine learning when training data contains labeling errors.\"},{\"question\":\"How do the authors reinterpret mislabeled detection methods?\",\"answer\":\"Most existing methods are described as probing trained machine learning models using a few shared underlying principles.\"},{\"question\":\"What does the benchmark evaluate and why is it important?\",\"answer\":\"It benchmarks methods under NCAR and NNAR labeling noise using imperfect labeling rules across multiple tasks, revealing both insights and limitations in realistic settings.\"}]","Mislabeled examples detection viewed as probing machine learning models - concepts, survey and extensive benchmark | PDF",1785736340,108,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mislabeled-examples-detection-viewed-as-probing-machine-learning-models-concepts-survey-and-extensive-benchmark","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mislabeled-examples-detection-viewed-as-probing-machine-learning-models-concepts-survey-and-extensive-benchmark/121581/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core problem addressed in this paper?","Question",{"text":75,"@type":76},"The paper studies how to detect mislabeled examples in supervised machine learning when training data contains labeling errors.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors reinterpret mislabeled detection methods?",{"text":80,"@type":76},"Most existing methods are described as probing trained machine learning models using a few shared underlying principles.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the benchmark evaluate and why is it important?",{"text":84,"@type":76},"It benchmarks methods under NCAR and NNAR labeling noise using imperfect labeling rules across multiple tasks, revealing both insights and limitations in realistic settings.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]