[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124134-en":3,"doc-seo-124134-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},124134,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Handling Imbalanced Datasets in Machine Learning - Challenges, Approaches, and Best Practices","Machine learning model performance depends not only on accuracy, but also on data quality and balance. When class distributions are imbalanced, a model can achieve high overall accuracy by favoring the majority class, producing biased and weak predictions for the minority class. Imbalanced datasets are commonly characterized by limited minority samples versus abundant majority samples, which degrades generalization. The study provides a comprehensive overview of practical strategies, including resampling, ensemble methods, and cost-sensitive learning.","Handling Imbalanced Datasets in Machine Learning: Challenges, Approaches, and Best Practices  \nRusmaAnieza Ruslan 1, NureizeArbaiy1  \n1 Faculty of Computer Science and Information Technology,  \nUniversiti Tun Hussein Onn Malaysia, Parit Raja, Batu Pahat, 86400, MALAYSIA  \n*[Corresponding Author: hi230027@student.uthm.edu.my](Corresponding Author: hi230027@student.uthm.edu.my)[ ](Corresponding Author: hi230027@student.uthm.edu.my)DOI: [https://doi.org/10.30880/jastec.2024.01.02.003](https://doi.org/10.30880/jastec.2024.01.02.003)  \nArticle Info  \nReceived: 2 August 2024  \nAccepted: 1 October 2024  \nAvailable online: 12 November 2024  \nKeywords  \nImbalanced dataset, machine learning, resampling, ensemble method, cost sensitive laerning  \nAbstract  \nDetermining the performance of a machine learning model is usually about the model's ability to make accurate predictions, which is assessed using an accuracy measure. However, other characteristics such as the quality and balance of the data must also be examined. Models may tend to make certain predictions that provide a high percentage of accurate predictions but have poor overall performance. There are balanced and imbalanced data situations in the dataset. Animbalanced dataset is a dataset that contains a minority class with a limited sample compared to the majority class. This makes it more likely that the model will favor the majority class, resulting in biased predictions and poor performance for the minority class. Therefore, it is important to remove the imbalance between the classes so that the model can make more accurate predictions. Several methods to solve this problem can be found in the literature, including the resampling method. Therefore, in this study, a comprehensive overview of techniques for dealing with imbalanced datasets in machine learning is given. This technique includes resampling techniques, ensemble methods and cost-sensitive learning. The study concludes that all techniques can be effective strategies for dealing with imbalanced class set problems and improving the performance of classification models in different domains.  \n1. Introduction  \nMachine learning (ML) is a subfield of artificial intelligence (AI) that focuses on the development of models and algorithms that can recognize patterns in data to make predictions and decisions [1, 3] without being explicitly programmed. There are different types of ML, including supervised, unsupervised, semi-supervised and reinforcement learning [2] . ML algorithms learn from data, and datasets play a crucial role in the development and performance of ML models. Datasets consist of rows, which represent observations, and columns, which represent features or characteristics. Datasets can be numeric, categorical, univariate, multivariate or time series and are essential for training, testing and evaluating ML models.  \nDatasets are a fundamental element of machine learning, as the quality and characteristics of the dataset significantly affect the performance and generalizability of ML models. Datasets can be labeled, i.e. each instance is associated with a corresponding label or target value, or unlabeled, i.e. the dataset has no predefined labels. Labeled datasets are used in supervised learning, while unlabeled datasets are used in unsupervised learning. Datasets can also be structured, unstructured, text-based or image-based and require different processing and analysis techniques. Each dataset has different characteristics in terms of size, complexity and balance.  \nThe balance of datasets is divided into two categories: balanced datasets and imbalanced datasets. In a balanced dataset, each class is equally represented whether positive or negative, while an imbalanced dataset hasan uneven distribution with underrepresented classes. The balance of the dataset balance is critical to the performance of each model in machine learning. Ifa minority class is significantly underrepresented compared toa majority cl","cbCaiaaD6maYZTST","https://ap.wps.com/l/cbCaiaaD6maYZTST","pdf",725997,1,"English","en",105,"# Introduction\n# Literature Review","[{\"question\":\"What makes a dataset “imbalanced” in machine learning?\",\"answer\":\"A dataset is imbalanced when the number of samples differs across target classes, typically with a minority class having far fewer examples than the majority class. This skew makes biased learning more likely.\"},{\"question\":\"Why does class imbalance harm model predictions?\",\"answer\":\"Imbalance can cause the model to favor the majority class because it is better represented, leading to inaccurate or biased predictions for the minority class. This is especially harmful in critical tasks like fraud detection and medical diagnosis.\"},{\"question\":\"What strategies are discussed to handle imbalanced datasets?\",\"answer\":\"The document highlights approaches including resampling (oversampling minority or undersampling majority), class weighting, ensemble methods, and cost-sensitive learning. These techniques aim to improve classification performance in imbalanced settings.\"}]","Handling Imbalanced Datasets in Machine Learning - Challenges, Approaches, and Best Practices | PDF",1785820635,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"handling-imbalanced-datasets-in-machine-learning-challenges-approaches-and-best-practices","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/handling-imbalanced-datasets-in-machine-learning-challenges-approaches-and-best-practices/124134/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What makes a dataset “imbalanced” in machine learning?","Question",{"text":74,"@type":75},"A dataset is imbalanced when the number of samples differs across target classes, typically with a minority class having far fewer examples than the majority class. This skew makes biased learning more likely.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Why does class imbalance harm model predictions?",{"text":79,"@type":75},"Imbalance can cause the model to favor the majority class because it is better represented, leading to inaccurate or biased predictions for the minority class. This is especially harmful in critical tasks like fraud detection and medical diagnosis.",{"name":81,"@type":72,"acceptedAnswer":82},"What strategies are discussed to handle imbalanced datasets?",{"text":83,"@type":75},"The document highlights approaches including resampling (oversampling minority or undersampling majority), class weighting, ensemble methods, and cost-sensitive learning. These techniques aim to improve classification performance in imbalanced settings.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]