[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128740-en":3,"doc-seo-128740-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128740,1099523885074,"Ivy","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","A Stacking Ensemble Machine Learning Strategy for COVID-19 Seroprevalence Estimations in the USA - based on Genetic Programming","The COVID-19 pandemic highlighted the need for reliable methods to study epidemic spread, since official infection prevalence data derived from PCR and antigen reports can be biased and inconsistent. A predictive modeling framework is presented using Genetic Programming-based stacking ensemble techniques to estimate SARS-CoV-2 seroprevalence in a target population from multiple prevalence sources, including official data, wastewater-derived estimates, and large survey-based estimates. Results show improved prediction accuracy compared with conventional stacking based on Linear Regression.","A Stacking Ensemble Machine Learning Strategy for COVID-19 Seroprevalence Estimations in the USA based on Genetic Programming  \nGontzal Sagastabeitia∗ , Josu Doncel∗ , Antonio Fernndez Anta†, Jose Aguilar†‡ and Juan Marcos Ramirez†  \n∗ University of the Basque Country UPV/EHU, Leioa, Biscay  \nEmail: [gontzal.sagastabeitia@ehu.eus](gontzal.sagastabeitia@ehu.eus)  \n†IMDEA Networks Institute, Madrid, Spain  \n‡CEMISID, University of the Andes, Merida, Venezuela  \nAbstract—The COVID-19 pandemic exposed the importance of research on the spread of epidemic diseases. In the case of COVID-19, official data about infection prevalence was based on PCR and antigen tests reports, which can be unreliable. In our work, we construct prediction models based on Genetic Programming to estimate the SARS-CoV-2 seroprevalence of a given population from multiple estimates of the COVID- 19 prevalence (official prevalence data, estimates derived from wastewater data, and estimates obtained from massive surveys with different rules and ML methods). To do that, we propose the use of stacking techniques based on Genetic Programming to obtain Machine Learning Ensemble Methods. Our approach produces more accurate prediction models than conventional stacking techniques based on Linear Regression.  \nI. INTRODUCTION  \nThe Coronavirus disease 2019 (COVID-19), caused by the severe acute respiratory syndrome Coronavirus 2 (SARS-CoV- 2) [1], has raised public interest in epidemics. During the pandemic, media outlets mainly reported daily updates on the number of COVID-19 infections, hospitalisations, and deaths to provide information about the spread of the disease. COVID-19 estimated number of cases was primarily obtained from large-scale screening using PCR and antigen tests [2] . However, this method may not be the most reliable source of information when attempting to understand the full scope of the pandemic and accurately determine the percentage of the population affected, since the accuracy of the information obtained from test screening is affected by various factors such as the limited availability of test kits [3] (especially at the beginning of the pandemic), the time between infection and the test timing [4], and the high number of asymptomatic infected individuals [5] .  \nThe traditional approach for estimating the proportion of previously infected individuals within a population relies on the measurement of seroprevalence. Specifically, seroprevalence refers to the proportion of individuals who test positive for a specific antibody in their blood [6] . In the case of COVID-19, a seropositive individual is a person who has SARS-CoV-2 antibodies in their blood. The presence of antibodies is considered sufficient evidence to confirm past infection, even without a positive test result. Multiple  \nseroprevalence studies were conducted during the COVID-19 pandemic in different countries, which required blood analyses of thousands of individuals along multiple rounds [7], [8] . These campaigns required substantial resources for logistical and organisational purposes.  \nOn the other hand, many approaches have been proposed during the COVID-19 pandemic that rely on data analysis and artificial intelligence to estimate the number of daily cases accurately [3], [9], [10] . These methods exploit the ability of online tools to track health indicators in almost real-time by collecting vast amounts of data from self-reported information. Estimating the seroprevalence from this data type is required to provide healthcare systems with a less expensive method for tracking the spread of diseases. In this context, it is necessary to analyse Ensemble Methods that allow combining the different estimation approaches (regardless of whether they are based on machine learning techniques or not) . In particular, we are interested in stacking techniques as an ensemble learning strategy, since it allows learning how to combine the estimates of numerous machine learning models ","cbCaig4YFyAQUcyu","https://ap.wps.com/l/cbCaig4YFyAQUcyu","pdf",1156992,1,9,"English","en",105,"# Abstract\n# Introduction\n## State of the Art","[{\"question\":\"Why are PCR and antigen test reports considered unreliable for COVID-19 prevalence estimation?\",\"answer\":\"Test-based accuracy can be affected by limited test kit availability, delays between infection and testing, and the presence of many asymptomatic infections.\"},{\"question\":\"What does the proposed method estimate and how is it defined?\",\"answer\":\"It estimates SARS-CoV-2 seroprevalence, defined as the proportion of individuals with detectable SARS-CoV-2 antibodies in their blood, indicating past infection.\"},{\"question\":\"How does the Genetic Programming stacking approach differ from conventional stacking?\",\"answer\":\"The method uses Genetic Programming to build stacking-based ensemble machine learning models and achieves more accurate predictions than conventional stacking approaches based on Linear Regression.\"}]","A Stacking Ensemble Machine Learning Strategy for COVID-19 Seroprevalence Estimations in the USA - based on Genetic Programming | PDF",1786003017,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-stacking-ensemble-machine-learning-strategy-for-covid-19-seroprevalence-estimations-in-the-usa-based-on-genetic-programming","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-stacking-ensemble-machine-learning-strategy-for-covid-19-seroprevalence-estimations-in-the-usa-based-on-genetic-programming/128740/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are PCR and antigen test reports considered unreliable for COVID-19 prevalence estimation?","Question",{"text":76,"@type":77},"Test-based accuracy can be affected by limited test kit availability, delays between infection and testing, and the presence of many asymptomatic infections.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What does the proposed method estimate and how is it defined?",{"text":81,"@type":77},"It estimates SARS-CoV-2 seroprevalence, defined as the proportion of individuals with detectable SARS-CoV-2 antibodies in their blood, indicating past infection.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the Genetic Programming stacking approach differ from conventional stacking?",{"text":85,"@type":77},"The method uses Genetic Programming to build stacking-based ensemble machine learning models and achieves more accurate predictions than conventional stacking approaches based on Linear Regression.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]