[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118765-en":3,"doc-seo-118765-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118765,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Shedding light on underrepresentation and Sampling Bias in machine learning","Accurately measuring discrimination is central to evaluating fairness in trained machine learning models, because measurement bias can amplify or hide existing disparities. The work distinguishes sources of bias by focusing on sampling-related effects that are not consistently defined in the literature. It introduces two variants—sample size bias and underrepresentation bias—then studies how discrimination decomposes into variance, bias, and noise. The paper also challenges the assumption that fairness can be improved merely by collecting more samples from underrepresented groups.","SHEDDING LIGHT ON UNDERREPRESENTATION AND SAMPLING  \nBIAS IN MACHINE LEARNING  \narXiv :2306 .05068v 1 [ cs .LG] 8 Jun 2023  \nSami Zhioua  \nINRIA, LIX, ´Ecole Polytechnique Palaiseau, Paris, France [zhioua@lix.polytechnique.fr](zhioua@lix.polytechnique.fr)  \nRta Binkyt  \nINRIA, LIX, ´Ecole Polytechnique Palaiseau, Paris, France [ruta.binkyte@inria.fr](ruta.binkyte@inria.fr)  \nABSTRACT  \nAccurately measuring discrimination is crucial to faithfully assessing fairness of trained machine learning (ML) models. Any bias in measuring discrimination leads to either amplification or underestimation of the existing disparity. Several sources of bias exist and it is assumed that bias resulting from machine learning is born equally by different groups (e.g. females vs males, whites vs blacks, etc.) . If, however, bias is born differently by different groups, it may exacerbate discrimination against specific sub-populations. Sampling bias, is inconsistently used in the literature to describe bias due to the sampling procedure. In this paper, we attempt to disambiguate this term by introducing clearly defined variants of sampling bias, namely, sample size bias (SSB) and underrepresentation bias (URB) . We show also how discrimination can be decomposed into variance, bias, and noise. Finally, we challenge the commonly accepted mitigation approach that discrimination can be addressed by collecting more samples of the underrepresented group.  \nKeywords ML Fairness ¨ Representation Bias ¨ Sampling Bias  \n1 Introduction  \nWith the ubiquitous use of machine learning (ML) systems to inform decisions with critical impacts on human lifes (e.g. job hiring, college admission, security screening), fairness is emerging as an important requirement for the safe use of these technologies. A failure to guarantee fairness may create or amplify discrimination against individuals or specific sub-populations (e.g. minority groups) . Such anomaly can initiate a vicious cycle that can be perpetuated and eventually resulting in severe consequences.  \nDiscrimination in ML decisions can originate from several types of bias as described in the literature. For instance, The Centre for Evidence-Based Medicine (CEBM) at the University of Oxford is maintaining a list of 62 different sources of bias [22] . More related to ML, Mehrabi et al. [20] classify the sources of bias into three categories depending on when the bias is introduced in the automated decision loop. For instance, measurement bias [14, 24] can be introduced at the data generation step and is a result of measuring a feature using a proxy variable instead of an ideal variable (e.g. using SAT score variable as a measure for the qualification feature) .  \nAnother, more common, category of bias occurs when the ML model is trained using a limited number of samples. This produces an inaccurate model and the inaccuracy will typically be born differently by different sub-populations which leads to a discrimination. Two famous examples of ML discrimination fall into this category of bias. The first is COMPAS software [7] used by several states in US to help predict whether a defendant will recidivate in the next two years if she is released. The software is found to be discriminatory against african-americans as the false positive rate (FPR) was higher for african-americans compared to other ethnicities, but the false negative rate (FNR) was lower [1] . The second example is related to face recognition technology (FRT) . Buolamwini et al. [3] found that several commercial FRT software have a significantly lower accuracy for individuals belonging to a specific sub-population, namely, dark-skinned females.  \nThis category of bias is inconsistently given various names in the literature (e.g. sampling bias, representation bias, data imbalance bias, etc.) and, to the best of our knowledge, is not formally defined. This paper is an attempt to  \nShedding light on underrepresentation and Sampling Bias in machine learning  \n","cbCaifoJXHmaYldN","https://ap.wps.com/l/cbCaifoJXHmaYldN","pdf",970095,1,17,"English","en",105,"# Introduction\n## Fairness as a requirement for ML decision-making\n## Types and timing of bias in automated decision loops\n## Sample-limited training and discrimination examples\n## Defining sampling-related bias variants\n## Metrics and decomposition of discrimination","[{\"question\":\"Why is precise measurement of discrimination important for ML fairness?\",\"answer\":\"If discrimination is measured with bias, it can either amplify or underestimate the existing disparity in trained models, leading to incorrect conclusions about fairness.\"},{\"question\":\"What are the two clarified variants of sampling bias proposed in the paper?\",\"answer\":\"The paper defines sample size bias (SSB) as bias from training with limited samples while keeping sub-populations in the same real proportions, and underrepresentation bias (URB) as bias from training with unequal sample counts across sub-populations.\"},{\"question\":\"How does the paper evaluate discrimination and its components?\",\"answer\":\"It uses multiple discrimination metrics such as differences in false positive rate, equal opportunity, ZOL, AUC, statistical disparity, and MSE for regression. For regression, it leverages prior work to decompose discrimination into noise, bias, and variance.\"}]","Shedding light on underrepresentation and Sampling Bias in machine learning | PDF",1785720119,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"shedding-light-on-underrepresentation-and-sampling-bias-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/shedding-light-on-underrepresentation-and-sampling-bias-in-machine-learning/118765/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is precise measurement of discrimination important for ML fairness?","Question",{"text":76,"@type":77},"If discrimination is measured with bias, it can either amplify or underestimate the existing disparity in trained models, leading to incorrect conclusions about fairness.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the two clarified variants of sampling bias proposed in the paper?",{"text":81,"@type":77},"The paper defines sample size bias (SSB) as bias from training with limited samples while keeping sub-populations in the same real proportions, and underrepresentation bias (URB) as bias from training with unequal sample counts across sub-populations.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the paper evaluate discrimination and its components?",{"text":85,"@type":77},"It uses multiple discrimination metrics such as differences in false positive rate, equal opportunity, ZOL, AUC, statistical disparity, and MSE for regression. For regression, it leverages prior work to decompose discrimination into noise, bias, and variance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]