[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119220-en":3,"doc-seo-119220-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119220,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Discriminative machine learning for maximal representative subsampling","Biased samples remain a common cause of distorted conclusions in social sciences and epidemiology. This work introduces two bias-mitigation methods based on positive-unlabeled learning that exploit auxiliary information from a representative dataset to estimate sample weights. The first, maximum representative subsampling (MRS), iteratively removes biased instances until the biased data become indistinguishable from the representative one. The second, Soft-MRS, adapts weights instead of discarding samples. Experiments on induced bias and a resilience-voting case study compare performance against existing techniques and guide when to use MRS versus Soft-MRS.","[www. nature.com/scientificreports](www. nature.com/scientificreports)  \nOPEN  \nDiscriminative machine learning for maximal representative subsampling  \nTony Hauptmann1*, Sophie Fellenz1, Laksan Nathan1, Oliver Tüscher2,3 & Stefan Kramer1  \nBiased population samples pose a prevalent problem in the social sciences. Therefore, we present two novel methods that are based on positive-unlabeled learning to mitigate bias. Both methods leverage auxiliary information from a representative data set and train machine learning classifiersto determine the sample weights. The first method, named maximum representative subsampling (MRS), uses a classifier to iteratively remove instances, by assigning a sample weight of 0, from the biased data set until it aligns with the representative one. The second method is a variant of MRS  \n– Soft-MRS – that iteratively adapts sample weights instead of removing samples completely. To assess the effectiveness of our approach, we induced artificial bias in a public census data set and examined the corrected estimates. We compare the performance of our methods against existing techniques, evaluating the ability of sample weights created with Soft-MRS or MRS to minimize differences and improve downstream classification tasks. Lastly, we demonstrate the applicability of the proposed methods in a real-world study of resilience research, exploring the influence of resilience on voting behavior. Through our work, we address the issue of bias in social science, amongst others, and provide a versatile methodology for bias reduction based on machine learning. Based on our experiments, we recommend to use MRS for downstream classification tasks and Soft-MRS for downstream tasks where the relative bias of the dependent variable is relevant.  \nA frequent challenge in social sciences and epidemiology is that a study sample is not representative for the entire population. Studies based on biased samples are not suitable for drawing conclusions because they lead to biased inferences about social processes1. Every phase of a research project, e.g. study design, data collection, or data analysis2,3 can induce bias, which is not dichotomous but can exist at various intensities. It is crucial for researchers to validate the presence of bias and in case of existence to mitigate is as much as possible to currectly analyze and interpret their findings.  \nExamples for bias are a survey that was conducted only in a specific city, although the goal was to generalize its results to the whole country or an online survey that depends on individual user participation, but people with a strong opinion on the topic have a higher likelihood to participate in the survey, leading to self-selection bias4. However, it is desirable to draw conclusions about all users or residents. If conclusions are drawn without bias correction, misleading descriptions of populations are produced and, ultimately, false conclusions are drawn5. Most bias cannot be completely removed without gathering more data, but bias reduction methods strive to mitigate bias as much as possible. Reducing bias can lead to more consistent findings and, therefore, saving time and resources, as fewer repeated studies are required6.  \nThis article proposes a bias reduction method based on positive and unlabeled training (PU learning7) in combination with auxiliary information from population data. It takes advantage of available representative population data sets, which have been rigorously collected to prevent potential bias, representing the target population with high precision. We call our method maximal representative subsampling (MRS), because the representative distribution is used to remove elements from the biased data set to create a maximal representative subsample, which resembles the representative distribution. Figure 1 shows the structure of our method.  \nIn more detail, the central aspect of MRS is to train random forests8 to differentiate instances from the ","cbCaiflEAsH66NQv","https://ap.wps.com/l/cbCaiflEAsH66NQv","pdf",2538922,1,13,"English","en",105,"# Background and problem of biased samples\n## Bias in study design, collection, and analysis\n## Consequences of uncorrected bias\n# Proposed methods\n## Maximum representative subsampling (MRS)\n## Soft-MRS\n# Method mechanism\n## Training a classifier to distinguish datasets\n## Iterative removal or reweighting\n# Evaluation and applications\n## Corrected estimates with induced bias\n## Comparison with existing techniques\n## Resilience research case study","[{\"question\":\"Why is bias correction important in social science and epidemiology studies?\",\"answer\":\"Biased samples are not representative of the full population, which leads to biased inferences and misleading conclusions. Bias can arise in study design, data collection, or data analysis, often at varying intensities.\"},{\"question\":\"How does maximum representative subsampling (MRS) reduce bias?\",\"answer\":\"MRS trains a classifier to distinguish instances from a biased dataset versus a representative dataset, then iteratively removes high-probability biased instances by assigning them weight 0 until the datasets become indistinguishable.\"},{\"question\":\"What is Soft-MRS and when is it preferred over MRS?\",\"answer\":\"Soft-MRS is a variant that iteratively adapts sample weights instead of fully removing samples. The work recommends MRS for downstream classification tasks and Soft-MRS when the relative bias of the dependent variable is especially relevant.\"}]","Discriminative machine learning for maximal representative subsampling | PDF",1785723156,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"discriminative-machine-learning-for-maximal-representative-subsampling","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/discriminative-machine-learning-for-maximal-representative-subsampling/119220/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is bias correction important in social science and epidemiology studies?","Question",{"text":76,"@type":77},"Biased samples are not representative of the full population, which leads to biased inferences and misleading conclusions. Bias can arise in study design, data collection, or data analysis, often at varying intensities.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does maximum representative subsampling (MRS) reduce bias?",{"text":81,"@type":77},"MRS trains a classifier to distinguish instances from a biased dataset versus a representative dataset, then iteratively removes high-probability biased instances by assigning them weight 0 until the datasets become indistinguishable.",{"name":83,"@type":74,"acceptedAnswer":84},"What is Soft-MRS and when is it preferred over MRS?",{"text":85,"@type":77},"Soft-MRS is a variant that iteratively adapts sample weights instead of fully removing samples. The work recommends MRS for downstream classification tasks and Soft-MRS when the relative bias of the dependent variable is especially relevant.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]