[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126409-en":3,"doc-seo-126409-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126409,962085571259,"Theodora","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Comparison of ANOVA and Chi-Square Feature Selection Methods to Improve Machine Learning Performance in Anemia Classification","Anemia is a common hematological disorder characterized by reduced blood hemoglobin levels, potentially causing serious complications when missed. Machine learning can support early diagnosis, but performance is often limited by irrelevant or excessive features. This study evaluates how ANOVA and Chi-Square feature selection affect three models—Naive Bayes, K-Nearest Neighbor, and Support Vector Machine—using a Kaggle dataset with 15,300 instances and 25 features. Metrics include accuracy, precision, recall, and F1-score before and after feature selection. Results show clear gains after reduction, with the SVM + ANOVA setup reaching 94.61% accuracy, while non-selected models remain below 90%, supporting statistical feature selection to improve generalization for medical classification.","Comparison of ANOVA and Chi-Square Feature Selection Methods to Improve Machine Learning Performance in Anemia Classification  \nTiko Nur Annisa*1, Jasmir Jasmir2, Nurhadi Nurhadi3  \n1Magister of Information Systems Student, Universitas Dinamika Bangsa, Indonesia  \n2,3Magister of Information Systems, Universitas Dinamika Bangsa, Indonesia  \n[Email:](Email:1tikonurannisa10@gmail.com)[1](Email:1tikonurannisa10@gmail.com)[tikonurannisa10@gmail.com](Email:1tikonurannisa10@gmail.com)  \nReceived : Jul 2, 2025; Revised : Aug 10, 2025; Accepted : Aug 12, 2025; Published : Aug 18, 2025  \nAbstract  \n\n| Anemia is a prevalent hematological condition marked by decreased hemoglobin concentration in the blood , which can lead to serious health complications if undetected. Although machine learning has shown potential in supporting early diagnosis, its effectiveness is often hindered by irrelevant or excessive features. This study investigates the impact of ANOVA and Chi-Square feature selection methods in improving the effectiveness of three distinct machine learning models algorithms, Naive Bayes, K-Nearest Neighbor (KNN), and Support Vector Machine (SVM) for anemia classification. Using a Kaggle dataset consisting of 15,300 instances and 25 features, the evaluation of each model was conducted with reference to its accuracy, precision, recall, and F1-score, both before and after applying feature selection. Experimental results show a substantial improvement in classification performance after featureselection, with the SVM + ANOVA combination achieving the highest accuracy of 94.61% . In contrast, models without feature selection performed below 90%, highlighting the need for appropriate feature reduction techniques. This study contributes a comparative analysis framework for medical data classification, emphasizing the role of statistical feature selection in optimizing model accuracy. Its novelty lies in demonstrating consistent performance improvement across algorithms using real-world anemia data and providing evidence that ANOVA and Chi-Square can significantly enhance model generalization in medical diagnostic contexts.\u003Cbr>Keywords : Anemia, Classification, Improvement, Machine Learning, Performance |\n| --- |\n| This work is an open access article and licensed under a Creative Commons Attribution-Non Commercial\u003Cbr>4.0 International License\u003Cbr> |\n\n1. INTRODUCTION  \nAnemia is a common medical condition marked by a reduced oxygen-carrying capacity of the blood, typically due to a low red blood cell count or abnormal hemoglobin [1] . Erythrocytes play a vital role in delivering oxygen to body tissues and facilitating the removal of carbon dioxide [2] . Clinically, anemia is diagnosed when in women a hemoglobin concentration that is considered low is less than 11 g/dL, while in men it is less than 12 g/dL [3] .  \nAccording to WHO, around 1.62 billion people suffer from anemia globally, including 43% of children under five and 300,000 infants [4][5] . In 2019, the global prevalence reached 22.8%[6] . In Indonesia, anemia affects 27.2% of girls and 20.3% of boys aged 15–24, making it a significant public health issue, especially among adolescent girls[7] .  \nAnemia is commonly diagnosed through complete blood count, serum ferritin, and hemoglobin electrophoresis tests [8] . However, these manual methods face challenges such as limited resources, data interpretation errors, and delayed diagnosis [9][10] .  \nThe advancement of information technology, especially machine learning, supports the use of intelligent systems to assist in disease diagnosis more quickly, accurately, and efficiently, particularly in areas with limited medical resources [11][12][13] . Machine learning has been successfully applied to disease classification with promising results [14] . Model performance is strongly influenced by data  \nquality and feature relevance [15] . Irrelevant features may reduce computational efficiency and prediction accuracy [16][17], making featu","cbCaitKhUrcbMgJ3","https://ap.wps.com/l/cbCaitKhUrcbMgJ3","pdf",857110,5,1,16,"English","en",105,"# Abstract\n# Introduction\n## Anemia background and diagnosis\n## Machine learning for disease classification\n## Feature selection importance\n## Related work and research gap","[{\"question\":\"Why is feature selection important for anemia classification models?\",\"answer\":\"Irrelevant or excessive features can reduce computational efficiency and prediction accuracy. Feature selection helps optimize model performance and generalization.\"},{\"question\":\"Which machine learning models were evaluated in the study?\",\"answer\":\"The study assessed Naive Bayes, K-Nearest Neighbor (KNN), and Support Vector Machine (SVM). Each model was tested with feature selection and without it.\"},{\"question\":\"What combination achieved the best classification performance?\",\"answer\":\"The SVM combined with ANOVA feature selection achieved the highest accuracy of 94.61% in the experiments.\"}]","Comparison of ANOVA and Chi-Square Feature Selection Methods to Improve Machine Learning Performance in Anemia Classification | PDF",1785904912,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"comparison-of-anova-and-chi-square-feature-selection-methods-to-improve-machine-learning-performance-in-anemia-classification","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/comparison-of-anova-and-chi-square-feature-selection-methods-to-improve-machine-learning-performance-in-anemia-classification/126409/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is feature selection important for anemia classification models?","Question",{"text":77,"@type":78},"Irrelevant or excessive features can reduce computational efficiency and prediction accuracy. Feature selection helps optimize model performance and generalization.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Which machine learning models were evaluated in the study?",{"text":82,"@type":78},"The study assessed Naive Bayes, K-Nearest Neighbor (KNN), and Support Vector Machine (SVM). Each model was tested with feature selection and without it.",{"name":84,"@type":75,"acceptedAnswer":85},"What combination achieved the best classification performance?",{"text":86,"@type":78},"The SVM combined with ANOVA feature selection achieved the highest accuracy of 94.61% in the experiments.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]