[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123960-en":3,"doc-seo-123960-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123960,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Development of Machine Learning Algorithms in Student Performance Classification Based on Online Learning Activities","Educational data mining increasingly supports timely evaluation of students’ academic achievements, but algorithm performance must align with dataset characteristics. This study develops machine learning models that classify student performance using online learning activities. During data cleaning, VarianceThreshold removes features consisting entirely of zeros. Feature selection and SMOTE are applied in preprocessing to manage class imbalance, while k-nearest neighbors, multi-layer perceptron, and logistic regression are trained with 3-fold cross-validation and optimized via GridSearchCV. Accuracy, precision, recall, and F1-score guide evaluation, with MLP and LR reaching 100% and KNN improving after tuning.","Development of machine learning algorithms in student performance classification based on online learning activities  \nMuhammad Aqif Hadi Alias, Mohd Azri Abdul Aziz, Najidah Hambali, Mohd Nasir Taib  \nSchool of Electrical Engineering, College of Engineering, Universiti Teknologi MARA (UiTM), Selangor, Malaysia  \nArticle history:  \nReceived Jun 6, 2024 Revised Jul 19, 2024 Accepted Aug 6, 2024  \nKeywords:  \nClassification algorithms Feature selection  \nK-nearest neighbors Logistic regression Multi-layer perceptron Student performance Synthetic minority oversampling technique  \nCorresponding Author:  \nThe field of educational data mining has gained significant traction for its pivotal role in assessing students' academic achievements. However, to ensure the compatibility of algorithms with the selected dataset, it is imperative for a comprehensive analysis of the algorithms to be done. This study delved into the development of machine learning algorithms utilizing students' online learning activities to effectively classify their academic performance. In the data cleaning stage, we employed VarianceThreshold for discarding features that have all zeros. Feature selection andoversampling techniques were integrated into the data preprocessing, using information gain to facilitate efficient feature selection and synthetic minority oversampling technique (SMOTE) to address class imbalance. In the classification phase, three supervised machine learning algorithms: k-nearest neighbors (KNN), multi-layer perceptron (MLP), and logistic regression (LR) were implemented, with 3-fold cross-validation to enhance robustness. Classifiers’ performance underwent refinement through hyperparameter tuning via GridSearchCV. Evaluation metrics, encompassing accuracy, precision, recall, and F1-score, were meticulously measured for each classifier. Notably, the study revealed that both MLP and LR achieved impeccable scores of 100% across all metrics, while KNN exhibited a noticeable performance boost after using hyperparameter tuning.  \nThis is an open access article under the CC BY-SA license.  \nMohd Azri Abdul Aziz  \nSchool of Electrical Engineering, College of Engineering, Universiti Teknologi MARA (UiTM) Selangor, Malaysia  \nEmail: [azriaziz@uitm.edu.my](azriaziz@uitm.edu.my)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nThe performance of students in educational institutions has garnered increasing attention in which a substantial number of institutions have recognized this as a pivotal determinant in enhancing both the overall quality of the institutions and the educational outcomes of their students [1]–[3] . Identifying at-risk students early in the course offers us the capacity to implement interventions and initiatives to improve their academic performance [4]–[10] . Consequently, in the pursuit of a deeper comprehension of the learning process and the environmental factors influencing it, the field of educational data mining has gained notable momentum. This discipline assumes a critical role in the classification of students' academic achievements [11], [12] . The application of artificial intelligence in education, particularly machine learning, has increased, with the technology expected to give effective approaches to enhance education in general in the near future [13] . Intelligent m-learning systems have recently gained traction as a method of offering more effective education and flexible learning that is tailored to each student's learning ability [14] . The early attempts to enable such systems, for creating tools to help students and learning in a conventional or online context, through the use of machine learning techniques focused on anticipating student achievement in terms of grades attained [15] .  \nDespite the importance of data preprocessing procedures, classification models must be well-developed to provide more accurate classification performance, considering the suitability of the algorithms with the selected dataset. Thu","cbCailpAKp04X6JB","https://ap.wps.com/l/cbCailpAKp04X6JB","pdf",525848,1,11,"English","en",105,"# Abstract\n# Introduction\n## Background and motivation\n## Role of educational data mining and machine learning\n# Methods and modeling approach\n## Data preprocessing and cleaning\n## Feature selection and oversampling\n## Classification algorithms and validation\n# Results and evaluation","[{\"question\":\"How does the study prepare the data before classification?\",\"answer\":\"It cleans the dataset using VarianceThreshold to discard features that contain all zeros. Then it applies feature selection and SMOTE during preprocessing to address class imbalance.\"},{\"question\":\"Which classification algorithms are used in the study?\",\"answer\":\"The study implements three supervised algorithms: k-nearest neighbors (KNN), multi-layer perceptron (MLP), and logistic regression (LR).\"},{\"question\":\"How are model performance and hyperparameters evaluated?\",\"answer\":\"Models are trained with 3-fold cross-validation for robustness, and hyperparameters are tuned using GridSearchCV. Performance is measured with accuracy, precision, recall, and F1-score.\"}]","Development of Machine Learning Algorithms in Student Performance Classification Based on Online Learning Activities | PDF",1785819446,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"development-of-machine-learning-algorithms-in-student-performance-classification-based-on-online-learning-activities","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/development-of-machine-learning-algorithms-in-student-performance-classification-based-on-online-learning-activities/123960/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the study prepare the data before classification?","Question",{"text":75,"@type":76},"It cleans the dataset using VarianceThreshold to discard features that contain all zeros. Then it applies feature selection and SMOTE during preprocessing to address class imbalance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which classification algorithms are used in the study?",{"text":80,"@type":76},"The study implements three supervised algorithms: k-nearest neighbors (KNN), multi-layer perceptron (MLP), and logistic regression (LR).",{"name":82,"@type":73,"acceptedAnswer":83},"How are model performance and hyperparameters evaluated?",{"text":84,"@type":76},"Models are trained with 3-fold cross-validation for robustness, and hyperparameters are tuned using GridSearchCV. Performance is measured with accuracy, precision, recall, and F1-score.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]