[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118279-en":3,"doc-seo-118279-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118279,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Bias Correction in Machine Learning-based Classification of Rare Events - Symposium on Data Science and Statistics (SDSS 2024)","Online platform businesses can be identified using web-scraped texts, but this classification is difficult because online platforms are rare. The work develops a machine learning text classification approach designed to minimize false positives in order to improve rare-event identification. The method incorporates Bayesian probability calibration and ensemble voting, which substantially reduces bias in prevalence estimates compared with uncalibrated probability outputs. The approach also cuts the need for extensive manual validation.","Symposium on Data Science and Statistics (SDSS 2024) Bias Correction in Machine Learning-based Classification of Rare Events  \nLuuk Gubbels Marco Puts Piet Daas  \nAbstract  \nOnline platform businesses can be identified by using web-scraped texts. This is a classification problem that combines elements of natural language processing and rare event detection. Because online platforms are rare, accurately identifying them with Machine Learning algorithms is challenging. Here, we describe the development of a Machine Learning-based text classification approach that reduces the number of false positives as much as possible. It greatly reduces the bias in the estimates obtained by using calibrated probabilities and ensembles.  \nKeywords: Calibration, Population, Ensembles  \n1 Introduction  \nObtaining reliable information from a small or rare subpopulation is a challenging topic for many researchers. Approaches commonly used to find rare or so-called hard-toidentify groups are a screening survey, network sampling, area sampling, or a combination (Snijkers et al. 2013). Examples of a rare subpopulation are online platforms. These are defined as:  \nA digital service that facilitates interactions between two or more distinct but interdependent sets of users (whether firms or individuals) who interact through the service via the Internet (OECD, 2019) .  \nTo obtain a complete overview of all online platforms, a Support Vector Machine (SVM) classification model was developed and applied to the entire population of businesses with a website in the Netherlands. Here, it was found that: i) online platforms compose of 0.22% of the total population of businesses with a website and ii) considerable manual checking was needed in this process (Daas et al., 2023). The latter was predominantly the result of the large number of false positives produced by the model. In this paper, we describe the development of a fully automated approach that aims to reduce the number of false positive online platforms detected as much as possible. This, consequently, seriously reduces the bias in the total estimate of the number of online platforms obtained and the manual checking needed.  \n2 Data and methods  \nAll websites assigned to businesses in the Business Register (BR) of Statistics Netherlands were studied. This resulted in a list of 960,588 (unique) URLs of which 629,284 could actually be scraped (Daas et al. ,  \n2023) . Web scraping was performed as described in Daas and van der Doef (2021) and up to a maximum of 200 pages was collected per URL. The texts were extracted and processed as described in Daas et al.(2023) . Per URL, the texts were combined.  \nFor online platform classification , a set of 569 online platforms (positives) were identified by experts. To this set, a random sample of 1328 non-platforms (negatives), from the scraped websites linked to the BR , were added. This resulted in an 1897-sized dataset with 30% platforms and 70% non-platforms. To ensure the independent evaluation of the findings , a 228-sized dataset was created, containing 69 positive and 159 negative cases, by sampling (and removing) them from the 1897-sized dataset. The resulting 1669 dataset was used for model development.  \nAll scripts were written in Python (v.3.7) and the Machine Learning algorithms implemented in scikitlearn (v.0.24) were used. The Bayesian calibration method of Puts and Daas (2021) was used to correct the intrinsic prevalence of probability-producing classification models. The code is available on GitHub (Puts, 2023) . Multiple algorithms were tested but in this paper, only the results of Logistic Regression based models are shown. This classification model provided results almost similar to that of the original SVM model but could be trained much faster (Gubbels, 2023) . Multiple models, up to 10, were trained on random resamples (bootstraps) of the 1669-sized training set. The bootstraps were randomly split into a 70% training and a 30% test set. A","cbCain7aSTfzbjT4","https://ap.wps.com/l/cbCain7aSTfzbjT4","pdf",129463,1,2,"English","en",105,"# Introduction\n## Rare subpopulations and online platforms\n## Overview of the proposed automated approach\n# Data and methods\n## Data sources and dataset construction\n## Modeling pipeline and ensemble training\n## Probability calibration\n# Results\n## Reducing false positives\n## Using model probabilities\n## Calibrating probabilities","[{\"question\":\"Why is classifying online platforms considered a rare-event problem?\",\"answer\":\"Online platforms constitute only a small fraction of all businesses with websites. Reliable identification from this rare subpopulation is therefore difficult and can easily lead to biased estimates and many false positives.\"},{\"question\":\"What problem occurs when using raw model probabilities from Logistic Regression?\",\"answer\":\"Using uncalibrated probabilities increases the estimated number of online platforms, which in turn increases bias. The probabilities rely on assumptions that are not satisfied in this setting, so they do not behave like actual probabilities.\"},{\"question\":\"How do ensembles and probability calibration reduce bias and false positives?\",\"answer\":\"The approach combines multiple Logistic Regression models trained on bootstrapped resamples and aggregates their weighted votes. Bayesian calibration corrects the intrinsic prevalence produced by probability-based classifiers, yielding much lower bias and fewer false positives than the original SVM-based process.\"}]","Bias Correction in Machine Learning-based Classification of Rare Events - Symposium on Data Science and Statistics (SDSS 2024) | PDF",1785682782,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"bias-correction-in-machine-learning-based-classification-of-rare-events-symposium-on-data-science-and-statistics-sdss-2024","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":21},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/bias-correction-in-machine-learning-based-classification-of-rare-events-symposium-on-data-science-and-statistics-sdss-2024/118279/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is classifying online platforms considered a rare-event problem?","Question",{"text":74,"@type":75},"Online platforms constitute only a small fraction of all businesses with websites. Reliable identification from this rare subpopulation is therefore difficult and can easily lead to biased estimates and many false positives.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What problem occurs when using raw model probabilities from Logistic Regression?",{"text":79,"@type":75},"Using uncalibrated probabilities increases the estimated number of online platforms, which in turn increases bias. The probabilities rely on assumptions that are not satisfied in this setting, so they do not behave like actual probabilities.",{"name":81,"@type":72,"acceptedAnswer":82},"How do ensembles and probability calibration reduce bias and false positives?",{"text":83,"@type":75},"The approach combines multiple Logistic Regression models trained on bootstrapped resamples and aggregates their weighted votes. Bayesian calibration corrects the intrinsic prevalence produced by probability-based classifiers, yielding much lower bias and fewer false positives than the original SVM-based process.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":29,"slug":136},19,"General","general"]