[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121297-en":3,"doc-seo-121297-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121297,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Comparative Analysis of Data Balancing Techniques for Machine Learning Classification on Imbalanced Student Perception Datasets - Research Findings","Class imbalance in machine learning classification can bias predictions toward majority classes and weaken minority-class detection. This study compares how four sampling approaches—SMOTE, SMOTE + Tomek Links, ADASYN, and SMOTE + ENN—work when combined with nine machine learning algorithms on an imbalanced sentiment dataset from Class XI students. The dataset contains 300 instances and 36 features spanning textual, demographic, and sentiment labels (Positive, Neutral, Negative). Performance is evaluated via accuracy, precision, recall, F1-score, and AUC-ROC across train-test splits.","Comparative Analysis of Data Balancing Techniques for Machine Learning Classification on Imbalanced Student Perception Datasets  \nAhmad Saekhu*1, Berlilana2, Dhanar Intan Surya Saputra3  \n1,2,3Magister of Computer Science, Universitas Amikom Purwokerto, Jawa Tengah, Indonesia  \n[Email:](Email:1ahmadsaekhu0920@gmail.com)[1](Email:1ahmadsaekhu0920@gmail.com)[ahmadsaekhu0920@gmail.com](Email:1ahmadsaekhu0920@gmail.com)  \nReceived : Jan 4, 2025; Revised : Jan 30, 2025; Accepted : Feb 5, 2025; Published : Apr 26, 2025  \nAbstract  \n\n| Class imbalance is a common challenge in machine learning classification tasks, often leading to biased predictions toward the majority class. This study evaluates the effectiveness of various machine learning algorithms combined with advanced data balancing techniques in addressing class imbalance in a dataset collected from Class XI students of SMK Ma'arif 1 Kebumen. The dataset, comprising 300 instances and 36 features, includes textual attributes, demographic information, and sentiment labels categorized as Positive, Neutral, and Negative. Preprocessing steps included text cleaning, target encoding, handling missing data, and vectorization. Four sampling techniques—SMOTE, SMOTE + Tomek Links, ADASYN, and SMOTE + ENN—were applied to the training data to create balanced datasets. Nine machine learning algorithms, including CatBoost, Extra Trees, Random Forest, Gradient Boosting, and others, were evaluated using four train-test splits (60:40, 70:30, 80:20, and 90:10) . Model performance was assessed using metrics such as accuracy, precision, recall, F1-score, and AUCROC. The results demonstrate that SMOTE + Tomek Links is the most effective balancing technique, achieving the highest accuracy when paired with ensemble algorithms like Extra Trees and Random Forest. CatBoost also delivered competitive performance, showcasing its adaptability in imbalanced scenarios. The 90:10 train-test split consistently yielded the best results, emphasizing the importance of adequate training data for model generalization. This study highlights the critical role of data balancing techniques and robust algorithms in optimizing classification performance for imbalanced datasets and provides a framework for future research in similar contexts.\u003Cbr>Keywords : Class imbalance, Classification performance, Ensemble models, Machine learning, SMOTE |\n| --- |\n| This work is an open access article and licensed under a Creative Commons Attribution-Non Commercial\u003Cbr>4.0 International License\u003Cbr> |\n\n1. INTRODUCTION  \nThe presence of class imbalance is a persistent challenge in machine learning classification tasks, where one or more classes significantly outnumber others [1,2] . This imbalance often leads to biased model predictions favoring the majority class, resulting in poor detection of minority class instances [3] . Addressing class imbalance is critical, especially in applications where misclassifying the minority class can have significant consequences, such as fraud detection [4], medical diagnosis [5], or sentiment analysis [6] . In this study, we focus on the classification of imbalanced sentiment data collected from Class XI students of SMK Ma'arif 1 Kebumen, encompassing textual and demographic features to analyze students' perceptions.  \nAddressing class imbalance in machine learning has been an active area of research, with numerous studies proposing strategies to mitigate its adverse effects on classification performance. Various data balancing techniques and algorithmic advancements have been explored to tackle this issue [7] . One of the most widely studied techniques is the Synthetic Minority Oversampling Technique (SMOTE), which generates synthetic samples for the minority class by interpolating  \nbetween existing samples [8] . Over the years, SMOTE has been enhanced with methods like SMOTE + Tomek Links, which removes noisy and overlapping samples [9], and SMOTE + Edited Nearest Neighbors (ENN), which filters out","cbCaitONkLRmoPfW","https://ap.wps.com/l/cbCaitONkLRmoPfW","pdf",491883,1,14,"English","en",105,"# Introduction\n## Data imbalance and its impact on classification\n## Data balancing techniques (SMOTE variants, ADASYN)\n## Role of machine learning algorithms (ensembles and boosting)\n## Evaluation metrics and train-test split effects","[{\"question\":\"What problem does the study address in machine learning classification?\",\"answer\":\"The study addresses class imbalance, where one class dominates and the model becomes biased toward the majority class, harming minority-class detection.\"},{\"question\":\"Which data balancing techniques are compared?\",\"answer\":\"The study compares SMOTE, SMOTE + Tomek Links, ADASYN, and SMOTE + ENN applied to the training data to create more balanced datasets.\"},{\"question\":\"Which evaluation metrics and train-test splits are used?\",\"answer\":\"Model performance is measured using accuracy, precision, recall, F1-score, and AUC-ROC, and results are tested across four train-test splits: 60:40, 70:30, 80:20, and 90:10.\"}]","Comparative Analysis of Data Balancing Techniques for Machine Learning Classification on Imbalanced Student Perception Datasets - Research Findings | PDF",1785734954,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"comparative-analysis-of-data-balancing-techniques-for-machine-learning-classification-on-imbalanced-student-perception-datasets-research-findings","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/comparative-analysis-of-data-balancing-techniques-for-machine-learning-classification-on-imbalanced-student-perception-datasets-research-findings/121297/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in machine learning classification?","Question",{"text":75,"@type":76},"The study addresses class imbalance, where one class dominates and the model becomes biased toward the majority class, harming minority-class detection.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which data balancing techniques are compared?",{"text":80,"@type":76},"The study compares SMOTE, SMOTE + Tomek Links, ADASYN, and SMOTE + ENN applied to the training data to create more balanced datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"Which evaluation metrics and train-test splits are used?",{"text":84,"@type":76},"Model performance is measured using accuracy, precision, recall, F1-score, and AUC-ROC, and results are tested across four train-test splits: 60:40, 70:30, 80:20, and 90:10.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]