[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121214-en":3,"doc-seo-121214-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121214,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Machine Learning Classifiers and Data Synthesis Techniques to Tackle with Highly Imbalanced COVID-19 Data - Research Article","The COVID-19 pandemic drove demand for fast, reliable diagnostic models. This study compares Random Forest, Logistic Regression, and Decision Tree classifiers trained on heavily imbalanced, preprocessed tabular datasets containing 5086 negative and 558 positive cases. To address class imbalance, two data synthesis methods, CTGAN and TVAE, are used to construct balanced datasets, and classifiers are evaluated on both original and synthesized data for fair comparison. Results show Random Forest achieves the highest accuracy of 98.83% on the CTGAN-balanced dataset.","Mashhad  \nTechnology Association of Iran  \nMachine Learning Classifiers and Data Synthesis Techniques to Tackle with Highly Imbalanced COVID-19 Data  \nResearch Article  \nAvaz Naghipour1  , Mohammad Reza Abbaszadeh Bavil Soflaei2, Mostafa Ghaderi-Zefrehei3  \nDOI: 10.22067/cke.2024.88940.1121  \nAbstract The COVID-19 pandemic has highlighted the urgent need for rapid and accurate diagnostic methods. In this study, we evaluate three machine learning models—Random Forest (RF), Logistic Regression (LR) and Decision Tree (DT)—for detecting COVID-19 trained on preprocessed imbalanced datasets. The dataset used in this study is heavily imbalanced, with 5086 negative and 558 positive cases, posing a significant challenge for effective model training. To this end, we demonstrate the capability of two advanced data synthesis algorithms, Conditional Tabular Generative Adversarial Network (CTGAN) and Tabular Variational Autoencoder (TVAE), in addressing the class imbalance inherent in the dataset. The classifierstrained on the original as well as the balanced datasets were evaluated for comparison. Our findings reveal that RF obtains the highest accuracy of 98.83% on the CTGAN-balanced dataset. In conclusion, our results verify the potential of coupling data synthesis with traditional machine learning for the diagnosis ofCOVID-19. We hope that we become a valuable contributor to the ongoing AI for pandemic.  \nKeywords COVID-19 Detection, Machine Learning, CTGAN, TVAE, Class Imbalance.  \n1. Introduction  \nIn late 2019, a pneumonia outbreak originated in Wuhan, China, which was subsequently identified as being caused by the SARS-CoV-2 virus by the World Health Organization (WHO) [1] . SARS-CoV-2 is an enveloped virus with a positive-sense, single-stranded RNA genome [2] . This virus primarily targets the human respiratory system and is highly transmissible through respiratory droplets from coughing, sneezing, and direct physical contact [3] . Additionally, it can spread via contact with contaminated surfaces, where the virus can persist for several days depending on environmental conditions [4] .  \nCommon symptoms of the infection include fever, dry cough, loss of taste and smell, sore throat, and muscle pain [2]. The pandemic has had widespread impacts, leading to the postponement of school examinations, closure of offices, and widespread layoffs [5] . Estimates from the World Health Organization (WHO) show that the full death toll associated directly or indirectly with the COVID-19 pandemic between 1 January 2020 and 31 December 2021 was approximately 14.9 million [6] . The recent surge in data science has empowered healthcare professionals by providing tools to analyze massive datasets of health information for disease detection. This progress is driven by various techniques like deep learning, data mining, and especially machine learning (ML) . However, a key challenge remains: selecting the most appropriate ML algorithms that can learn effectively from existing data and make accurate predictions for entirely new cases [2] . ML is crucial in the healthcare sector, particularly for diagnosing diseases, detecting outbreaks, and preventing illnesses. ML algorithms are employed for numerous purposes, including predicting diabetes [7], forecasting the progression of Alzheimer's disease [8], heart disease [9], and other medical conditions. Due to the scarcity of tabular data on COVID-19, we tested our hypothesis using a dataset available on Kaggle (at this link), which clearly represents the clinical symptoms of COVID-19. This dataset, like many others in the field of ML, is heavily imbalanced, containing 5086 negative cases and 558 positive cases, resulting in a 1:9 ratio. Training ML models on imbalanced datasets poses several challenges: the models tend to be biased towards the majority class, leading to poor performance in detecting the minority class [10] . This imbalance can result in lower recall for the minority class, skewed accuracy m","cbCailbfbgQm8ZPP","https://ap.wps.com/l/cbCailbfbgQm8ZPP","pdf",802285,1,14,"English","en",105,"# Introduction\n## Dataset imbalance and training challenges\n## Study objective and approach\n## Overview of CTGAN and TVAE\n## Study structure","[{\"question\":\"Which machine learning models are evaluated for COVID-19 detection in the study?\",\"answer\":\"The study evaluates Random Forest, Logistic Regression, and Decision Tree for classifying COVID-19 based on preprocessed imbalanced tabular data.\"},{\"question\":\"How does the study address class imbalance in the COVID-19 dataset?\",\"answer\":\"It uses data synthesis to balance classes, specifically Conditional Tabular Generative Adversarial Network (CTGAN) and Tabular Variational Autoencoder (TVAE).\"},{\"question\":\"What performance result stands out from the experiments?\",\"answer\":\"Random Forest achieves the highest accuracy of 98.83% on the CTGAN-balanced dataset, supporting the value of combining data synthesis with traditional machine learning.\"}]","Machine Learning Classifiers and Data Synthesis Techniques to Tackle with Highly Imbalanced COVID-19 Data - Research Article | PDF",1785734384,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-classifiers-and-data-synthesis-techniques-to-tackle-with-highly-imbalanced-covid-19-data-research-article","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-classifiers-and-data-synthesis-techniques-to-tackle-with-highly-imbalanced-covid-19-data-research-article/121214/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning models are evaluated for COVID-19 detection in the study?","Question",{"text":75,"@type":76},"The study evaluates Random Forest, Logistic Regression, and Decision Tree for classifying COVID-19 based on preprocessed imbalanced tabular data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study address class imbalance in the COVID-19 dataset?",{"text":80,"@type":76},"It uses data synthesis to balance classes, specifically Conditional Tabular Generative Adversarial Network (CTGAN) and Tabular Variational Autoencoder (TVAE).",{"name":82,"@type":73,"acceptedAnswer":83},"What performance result stands out from the experiments?",{"text":84,"@type":76},"Random Forest achieves the highest accuracy of 98.83% on the CTGAN-balanced dataset, supporting the value of combining data synthesis with traditional machine learning.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]