[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119116-en":3,"doc-seo-119116-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},119116,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Supervised and Unsupervised Machine Learning Algorithms - An Empirical Evaluation","This study evaluates supervised and unsupervised machine learning algorithms across multiple real-world scenarios to determine which approaches perform best under different testing conditions. A pipeline-based workflow is used, including missing-value checking and systematic data preprocessing such as cleaning, integration, transformation, and reduction. The algorithms—KNN, Decision Trees, SVM, K-Means, and Agglomerative Hierarchical Clustering—are tested on medical, finance, ornithology, and outdoor image datasets. Comparative analysis highlights key strengths and weaknesses and reports accuracy-oriented results tied to sector-specific settings.","1  \nSupervised and Unsupervised Machine Learning Algorithms: An Empirical Evaluation  \nRajarajan Rajkumar, Li Zhang† and Vivian Sedov Department of Computer Science, Royal Holloway, University of London Surrey, TW20 0EX, UK  \n†E-mail: [li.zhang@rhul.ac.uk](li.zhang@rhul.ac.uk)  \nKamlesh Mistry  \nDepartment of Computer and Information Sciences, Northumbria University Newcastle, NE1 8ST, UK  \nMachine Learning (ML) algorithms are a subset of Artificial Intelligence that are applied to data with a primary focus of improving its accuracy over time by replicating and imitating the learning styles of human beings. Within this framework, several supervised and unsupervised learning algorithms are studied through different scenarios. The advantages and disadvantages of these algorithms are analyzed through these case studies.  \nKeywords: Supervised and unsupervised learning algorithms, Classification, Clustering.  \n1. Introduction  \nThe aim of this study is to understand which machine learning (ML) algorithm is more effective under different testing scenarios. ML is the use and development of computer systems that are able to learn and adapt without following explicit instructions, by using algorithms and statistical models to analyse and draw inferences from patterns in data. They have become popular due to their capability in executing complex tasks that are difficult to program and the ability to adapt to changes in circumstances through self-learning. Hence, it is crucial to differentiate and understand between algorithms that can be utilised to enhance data processing efficiency in our daily routines. This research aims to exploit distinctive learning behaviors of several supervised and unsupervised algorithms when tackling different classification/clustering tasks. The algorithms in the scope of this work are: Supervised Learning -K-nearest neighbour (KNN), Decision Tree learning and Support Vector Machines (SVM), and Unsupervised Learning-K-Means and Agglomerative Hierarchical  \n2  \nClustering (AHC) . These algorithms were applied and tested on four different datasets from various sectors: Medical, Finance, Ornithology and Outdoor Image datasets.  \n2. Related Work  \nThe study of [1] compared various supervised machine learning classification methods using a Diabetes data set with 786 instances and eight attributes as independent variables. Seven different algorithms were tested and SVM was found to be the most effective, followed by Naive Bayes and Random Forest. Sathya and Abraham [2] presented a comparative study of unsupervised and supervised learning models and their classification performance on higher education scenarios. Kohonen map of unsupervised learning model offers the most efficient solutions. Khanam [3] discussed the use of machine learning algorithms such as neural networks for early detection of diabetes. The Pima Indian Diabetes dataset was used, which contains information on 768 patients and nine attributes. Seven machine learning algorithms were studied and the combination of logistic regression and SVM was found to be the most effective for diabetes prediction. A neural network model with two hidden layers provided an accuracy rate of 88.6% . Perols [4] compared the performance of six different machine learning models in detecting financial statement fraud with different misclassification costs and fraud ratios. Logistic regression and SVMs were found to be the most effective methods, and only six predictors were consistently used across the algorithms. Abas [5] examined different data clustering algorithms, including K-Means, hierarchical clustering, selforganising maps, and expectation maximisation. The algorithms were evaluated based on factors such as dataset size, number of clusters, and dataset type.  \n3. The Proposed Methodologies  \nFor the model to give an effective output within reasonable performance, a pipeline must be created. The machine learning pipeline is an end-to-end framework which manages ","cbCaiu7TqZcfN7u3","https://ap.wps.com/l/cbCaiu7TqZcfN7u3","pdf",753339,1,"English","en",105,"# Introduction\n# Related Work\n# The Proposed Methodologies\n## Dataset Description\n## Data Pre-processing\n# Evaluation of Supervised Learning Algorithms\n## K-Nearest Neighbour\n## Decision Tree Learning\n## Support Vector Machines","[{\"question\":\"Which machine learning algorithms are evaluated in this study?\",\"answer\":\"The study evaluates supervised algorithms including K-nearest neighbour (KNN), Decision Tree learning, and Support Vector Machines (SVM), and unsupervised algorithms including K-Means and Agglomerative Hierarchical Clustering (AHC).\"},{\"question\":\"How are the datasets prepared before model evaluation?\",\"answer\":\"A pipeline is used with missing value checking and preprocessing steps covering data cleaning, integration, transformation, and reduction. These steps include handling missing values, smoothing noise, removing outliers, normalization, attribute selection, and aggregation.\"},{\"question\":\"What types of datasets are used for testing the algorithms?\",\"answer\":\"Four datasets are used: Medical, Finance, Ornithology, and Outdoor Image datasets. They target tasks such as heart attack prediction, financial-related analysis, bird species classification, and outdoor image interpretation.\"}]","Supervised and Unsupervised Machine Learning Algorithms - An Empirical Evaluation | PDF",1785722459,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"supervised-and-unsupervised-machine-learning-algorithms-an-empirical-evaluation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/supervised-and-unsupervised-machine-learning-algorithms-an-empirical-evaluation/119116/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning algorithms are evaluated in this study?","Question",{"text":75,"@type":76},"The study evaluates supervised algorithms including K-nearest neighbour (KNN), Decision Tree learning, and Support Vector Machines (SVM), and unsupervised algorithms including K-Means and Agglomerative Hierarchical Clustering (AHC).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the datasets prepared before model evaluation?",{"text":80,"@type":76},"A pipeline is used with missing value checking and preprocessing steps covering data cleaning, integration, transformation, and reduction. These steps include handling missing values, smoothing noise, removing outliers, normalization, attribute selection, and aggregation.",{"name":82,"@type":73,"acceptedAnswer":83},"What types of datasets are used for testing the algorithms?",{"text":84,"@type":76},"Four datasets are used: Medical, Finance, Ornithology, and Outdoor Image datasets. They target tasks such as heart attack prediction, financial-related analysis, bird species classification, and outdoor image interpretation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":28,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":28,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]