[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125137-en":3,"doc-seo-125137-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125137,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Mitigating Label Bias in Machine Learning - Fairness through Confident Learning - Paper","Discrimination arises when potentially biased agents overwrite underlying unbiased labels, producing biased datasets that unfairly harm specific groups and lead classifiers to reproduce these biases. This paper shows that, even with access only to biased labels, bias can be reduced by filtering the fairest instances within confident learning. Low self-confidence often signals label errors, but not always, especially for underrepresented groups, motivating a truncation of confidence scores, an expanded confidence interval, and co-teaching selection.","Mitigating Label Bias in Machine Learning: Fairness through Confident Learning  \nYixuan Zhang 1 , Boyu Li2 , Zenan Ling3 , Feng Zhou4 *  \n1 China-Austria Belt and Road Joint Laboratory on AI and AM, Hangzhou Dianzi University, China  \n2Data Science Institute, University of Technology Sydney, Australia  \n3 School of Electronic Information and Communications, Huazhong University of Science and Technology, China  \n4 Center for Applied Statistics and School of Statistics, Renmin University of China, China  \n[yixuan.zhang@hdu.edu.cn](yixuan.zhang@hdu.edu.cn), [boyu.li@student.uts.edu.au](boyu.li@student.uts.edu.au), [lingzenan@hust.edu.cn](lingzenan@hust.edu.cn), [feng.zhou@ruc.edu.cn](feng.zhou@ruc.edu.cn)  \nAbstract  \nDiscrimination can occur when the underlying unbiased labels are overwritten by an agent with potential bias, resulting in biased datasets that unfairly harm specific groups and cause classifiers to inherit these biases. In this paper, we demonstrate that despite only having access to the biased labels, it is possible to eliminate bias by filtering the fairest instances within the framework of confident learning. In the context of confident learning, low self-confidence usually indicates potential label errors; however, this is not always the case. Instances, particularly those from underrepresented groups, might exhibit low confidence scores for reasons other than labeling errors. To address this limitation, our approach employs truncation of the confidence score and extends the confidence interval of the probabilistic threshold. Additionally, we incorporate with co-teaching paradigm for providing amore robust and reliable selection of fair instances and effectively mitigating the adverse effects of biased labels. Through extensive experimentation and evaluation of various datasets, we demonstrate the efficacy of our approach in promoting fairness and reducing the impact of label bias in machine learning models.  \nIntroduction  \nIn recent decades, we have observed a shift towards an AIdriven society, where machine learning techniques have profoundly affected various aspects of our lives, including finance (Khandani, Kim, and Lo 2010), recruitment (Faliagka et al. 2012) and law (Dressel and Farid 2018) . When algorithmic decisions have a significant impact on our lives, it is crucial for decision-makers or regulators to have confidence in the algorithm’s performance. While deep neural networks possess the capability to learn complex patterns from input data, they encounter a fundamental challenge in the data they learn from: the labels can be influenced by sensitive information, causing the neural networks to capture and reproduce these undesirable associations. Therefore, to ensure fairness in decision-making, it becomes essential to mitigate the impact of such undesirable relationships.  \nResearch on training fair machine learning models has then received a lot of attention. To achieve fairness, one  \n* Corresponding author.  \nCopyright © 2024, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \ncan incorporate fairness constraints into the learning objectives (Bilal Zafar et al. 2015; Cotter, Jiang, and Sridharan 2019; Donini et al. 2018; Rezaei et al. 2020; Roh et al. 2020), or modify the model’s predictions using threshold adjustments and calibration to align them with fairness constraints (Hardt, Price, and Srebro 2016; Kim, Ghorbani, and Zou 2019; Lohia et al. 2019; Petersen et al. 2021), or use data manipulation or representation learning methods before model training (Calmon et al. 2017; Choi et al. 2020; Jiang and Nachum 2020; Kamiran and Calders 2012) . Nonetheless, the aforementioned methods primarily focus on modifying the machine learning model to tackle the bias problem. There has been limited effort, as indicated by Jiang and Nachum (2020), in directly addressing the biased data itself, despite the fact that often the training data itself ","cbCaiktk1ruwlHkQ","https://ap.wps.com/l/cbCaiktk1ruwlHkQ","pdf",240905,1,9,"English","en",105,"# Abstract\n# Introduction\n## Fairness challenges in decision-making\n## Prior approaches to training fair models\n## Directly addressing biased data vs. model changes\n## Data selection under confident learning","[{\"question\":\"How does label bias create unfair outcomes in machine learning?\",\"answer\":\"When unbiased ground-truth labels are overwritten by a potentially biased agent, the resulting biased dataset harms specific groups. Classifiers then inherit and reproduce these undesirable associations.\"},{\"question\":\"Why can low self-confidence scores be misleading?\",\"answer\":\"Low confidence may indicate label errors, but in imbalanced or underrepresented settings, instances can receive low confidence for reasons other than incorrect labeling.\"},{\"question\":\"What does the proposed approach do to mitigate label bias using confident learning?\",\"answer\":\"It truncates the confidence scores, extends the probabilistic threshold’s confidence interval, and uses a co-teaching paradigm to select more robust and reliable fair instances for retraining.\"}]","Mitigating Label Bias in Machine Learning - Fairness through Confident Learning - Paper | PDF",1785896861,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mitigating-label-bias-in-machine-learning-fairness-through-confident-learning-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mitigating-label-bias-in-machine-learning-fairness-through-confident-learning-paper/125137/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does label bias create unfair outcomes in machine learning?","Question",{"text":75,"@type":76},"When unbiased ground-truth labels are overwritten by a potentially biased agent, the resulting biased dataset harms specific groups. Classifiers then inherit and reproduce these undesirable associations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why can low self-confidence scores be misleading?",{"text":80,"@type":76},"Low confidence may indicate label errors, but in imbalanced or underrepresented settings, instances can receive low confidence for reasons other than incorrect labeling.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the proposed approach do to mitigate label bias using confident learning?",{"text":84,"@type":76},"It truncates the confidence scores, extends the probabilistic threshold’s confidence interval, and uses a co-teaching paradigm to select more robust and reliable fair instances for retraining.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]