[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124677-en":3,"doc-seo-124677-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124677,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","COVID-19 Symptoms Clustering and Severity Classification Using Machine Learning Approach - Research Summary","COVID-19 is highly contagious and evolves through new variants, making symptom identification critical for limiting transmission. This study examines a Kaggle dataset derived from surveys across participants from at least ten countries and labels symptoms into four severity levels aligned with WHO guidance and Indian Ministry recommendations. Supervised learning tests seven classifiers, while unsupervised learning applies K-Means and Expectation-Maximization clustering. Results show weak classification and clustering performance, raising concerns about dataset validity and reliability.","COVID-19: Symptoms Clustering and Severity Classification Using Machine Learning Approach  \nNurul Fathia Mohamand Noor 1, Herold Sylvestro Sipail1, Norulhusna Ahmad1, Bayram Annanurov2, Norliza Mohd Noor1 *  \n1Razak Faculty of Technology and Informatics,  \nUniversity Teknologi Malaysia, Kuala Lumpur, 54100, MALAYSIA  \n2Center for Biomedical Informatics.  \nWake Forest School of Medicine, Winston-Salem, NC, USA  \n*Corresponding Author  \nDOI: [https://doi.org/10.30880/ijie.2023.15.03.001](https://doi.org/10.30880/ijie.2023.15.03.001)  \nReceived 30 October 2022; Accepted 29 December 2022; Available online 31 July 2023  \nAbstract: COVID-19 is an extremely contagious illness that causes illnesses varying from either the common cold to more chronic illnesses or even death. The constant mutation of a new variant of COVID-19 makes it important to identify the symptom of COVID-19 in order to contain the infection. The use of clustering and classification in machine learning is in mainstream use in different aspects of research, especially in recent years to generate useful knowledge on COVID-19 outbreak. Many researchers have shared their COVID-19 data on public database and alot of studies have been carried out. However, the merit of the dataset is unknown and analysis need to be carried by the researchers to check on its reliability. The dataset that is used in this work was sourced from the Kaggle website. The data was obtained through a survey collected from participants of various gender and age who had been to at least ten countries. There are four levels of severity based on the COVID-19 symptom, which was developed in accordance to World Health Organization (WHO) and the Indian Ministry of Health and Family Welfare recommendations. This paper presented an inquiry on the dataset utilising supervised and unsupervised machine learning approaches in order to better comprehend the dataset. In this study, the analysis of the severity group based on the COVID-19 symptoms using supervised learning techniques employed a total of seven classifiers, namely the K-NN, Linear SVM, Naive Bayes, Decision Tree (J48), Ada Boost, Bagging, and Stacking. For the unsupervised learning techniques, the clustering algorithm utilized in this work are Simple K-Means and ExpectationMaximization. From the result obtained from both supervised and unsupervised learning techniques, we observed that the result analysis yielded relatively poor classification and clustering results. The findings for the dataset analysed in this study do not appear to be providing the correct result for the symptoms categorized against the severity level which raises concerns about the validity and reliability of the dataset.  \nKeywords: COVID-19 symptom, machine learning, classification  \n1. Introduction  \nThe COVID-19 outbreak has swept over the world since its onset in November 2019, infecting more than a hundred million individuals and killing more than 4 million people as of July 8, 2021 [1] . COVID-19 vaccination program may have been deployed globally since December 2020, with efficacy varying from 50.38 per cent to 95 per cent. While immunisation is the most effective strategy to contain the virus, there is concern that COVID-19 mutations can render the vaccination ineffective [2] . Immunisation will be a lengthy process as various COVID-19 variations evolve. Therefore, rapid COVID-19 detection alternatives are essential to minimise the virus from spreading.  \nAt the end of 2019, a pneumonia outbreak with an unknown aetiology was identified at the Wuhan Market, located in the province of Hubei, China. [3] . In January 2020, the unknown virus that caused it was identified by the World Health Organization (WHO) as COVID-19. This novel coronavirus is called severe acute respiratory syndrome coronavirus, SARS-CoV-2 [4] . The WHO declared a pandemic in February 2020 due to the epidemic and confirmed cases worldwide. [5]. More than a year later, in July 2021, the disease affecte","cbCaibk8UnCmTM6G","https://ap.wps.com/l/cbCaibk8UnCmTM6G","pdf",767192,1,14,"English","en",105,"# Introduction\n## Background on COVID-19 outbreak and transmission\n## Symptom presentation and clinical recommendations\n## Machine learning approaches (supervised and unsupervised)","[{\"question\":\"What dataset and severity framework are used in the study?\",\"answer\":\"The work uses a Kaggle-sourced survey dataset, and assigns symptoms into four severity levels based on WHO and Indian Ministry of Health and Family Welfare recommendations.\"},{\"question\":\"Which algorithms are applied for supervised and unsupervised learning?\",\"answer\":\"Seven supervised classifiers are evaluated (K-NN, Linear SVM, Naive Bayes, J48, Ada Boost, Bagging, Stacking). For unsupervised learning, Simple K-Means and Expectation-Maximization are used for clustering.\"},{\"question\":\"What do the findings suggest about the dataset and model performance?\",\"answer\":\"The study observes relatively poor classification and clustering results, where symptom categories do not align well with severity levels, indicating concerns about the dataset’s validity and reliability.\"}]","COVID-19 Symptoms Clustering and Severity Classification Using Machine Learning Approach - Research Summary | PDF",1785893860,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"covid-19-symptoms-clustering-and-severity-classification-using-machine-learning-approach-research-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/covid-19-symptoms-clustering-and-severity-classification-using-machine-learning-approach-research-summary/124677/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What dataset and severity framework are used in the study?","Question",{"text":75,"@type":76},"The work uses a Kaggle-sourced survey dataset, and assigns symptoms into four severity levels based on WHO and Indian Ministry of Health and Family Welfare recommendations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which algorithms are applied for supervised and unsupervised learning?",{"text":80,"@type":76},"Seven supervised classifiers are evaluated (K-NN, Linear SVM, Naive Bayes, J48, Ada Boost, Bagging, Stacking). For unsupervised learning, Simple K-Means and Expectation-Maximization are used for clustering.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the findings suggest about the dataset and model performance?",{"text":84,"@type":76},"The study observes relatively poor classification and clustering results, where symptom categories do not align well with severity levels, indicating concerns about the dataset’s validity and reliability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]