[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124149-en":3,"doc-seo-124149-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124149,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","A feature selection and scoring scheme for dimensionality reduction in a machine learning task","Feature selection is central to machine learning problems with high-dimensional datasets and large numbers of features, as it reduces dimensionality while improving predictive performance. Existing feature selection techniques often restrict dataset types and selection assumptions. This study proposes a generic feature selection approach using a statistical lift measure for binary classification, identifying an important feature subset and outperforming Chi-Square, Pearson Correlation, and Information Gain. Tests on lung cancer and happiness classification datasets evaluate multiple models using accuracy, precision, recall, and F1-score.","J. Nig. Soc. Phys. Sci. 7 (2025) 2273  \nA feature selection and scoring scheme for dimensionality reduction in a machine learning task  \nPhilemon Uten Emmoha,∗, Christopher Ifeanyi Ekeb , Timothy Mosesb  \na Department of Computer Science, Federal University Wukari, P.M.B 1020, Katsina-Ala Road, Wukari, Taraba State, Nigeria b Department of Computer Science, Federal University of Lafia, P.M. B 146, Lafia, Nasarawa State, Nigeria  \nAbstract  \nThe selection of important features is very vital in machine learning tasks involving high-dimensional dataset with large features. It helps to reduce the dimensionality of a dataset and improve model performance. Most of the feature selection techniques have restrictions on the kind of dataset to be used. This study proposed a feature selection technique based on statistical lift measure to select important features from a dataset. The proposed technique is a generic approach that can be used in any binary classification dataset problem. The technique successfully determined the most important feature subset and outperformed the existing techniques. The proposed technique was tested on lungs cancer dataset and happiness classification dataset. The effectiveness of the proposed technique in selecting important features subset was evaluated and compared with other existing techniques, namely Chi-Square, Pearson Correlation and Information Gain. The proposed and the existing techniques were evaluated on five machine learning models using four standard evaluation metrics such as accuracy, precision, recall and F1-score. The experimental results of the proposed technique on lung cancer dataset shows that logistic regression, decision tree, adaboost, gradient boost and random forest produced a predictive accuracy of 0.919%, 0.935%, 0.919%, 0.935% and 0.935% respectively, and that of happiness classification dataset produced a predictive accuracy of 0.758%, 0.689%, 0.724%, 0.655% and 0.689% on random forest, k-nearest neighbor, decision tree, gradient boost and cat boost respectively, which outperformed the existing techniques.  \nDOI:10.46481/jnsps.2025.2273  \nKeywords: Algorithm, Dataset, Dimensionality reduction, Feature selection  \nArticle History :  \nReceived: 26 July 2024  \nReceived in revised form: 04 November 2024  \nAccepted for publication: 05 November 2024  \nPublished: 14 December 2024  \n© 2025 The Author(s) . Published by the Nigerian Society of Physical Sciences under the terms of the Creative Commons Attribution 4.0 International license. Further distribution of this work must maintain attribution to the author(s) and the published article’s title, journal citation, and DOI.  \nCommunicated by: O. Akande  \n1. Introduction  \nMachine learning researchers and engineers face significant challenges when processing high-dimensional data. In highdimensional data, there are many features to be detected. There may be some unnecessary and unimportant features [1] . Several techniques have been developed to address the problem of  \n∗ Corresponding author: Tel.: +234-803-520-2835 .  \nEmail address: [philemon@fuwukari.edu.ng](philemon@fuwukari.edu.ng) (Philemon Uten Emmoh)  \nreducing irrelevant variables when performing data mining, machine learning, and other modelling tasks. Feature selection is a method used to select subsets of original features by eliminating unimportant or redundant features while maintaining the original qualities of the features that aid in visualizing and comprehending [2] . According to Peng et al. [3], feature selection aids in data comprehension, minimizes the need for computation, mitigates the consequences of the dimensionality curse, and enhances the predictive capabilities of models.  \nThe purpose of feature selection is to come up with sub-  \nEmmoh et al. / J. Nig. Soc. Phys. Sci. 7 (2025) 2273 2  \nsets of features from the input features that sufficiently represent the feature space [4] . According to Cherrington et al. [5], the choice of whether to keep important","cbCairG1xbfDwBY7","https://ap.wps.com/l/cbCairG1xbfDwBY7","pdf",473776,1,12,"English","en",105,"# Introduction\n## Feature selection overview\n## Need for dimensionality reduction in high-dimensional data\n## Feature selection categories (filtering, wrapping, embedded)","[{\"question\":\"What problem does the paper address in machine learning tasks?\",\"answer\":\"The paper addresses challenges of handling high-dimensional datasets with many features, including noisy, redundant, and uninformative variables that can degrade learning performance.\"},{\"question\":\"How does the proposed method select important features?\",\"answer\":\"It uses a statistical lift measure to score and select an important feature subset, designed as a generic approach for binary classification datasets.\"},{\"question\":\"Which baselines and evaluation metrics are used to compare techniques?\",\"answer\":\"The proposed technique is compared with Chi-Square, Pearson Correlation, and Information Gain, and evaluated across five machine learning models using accuracy, precision, recall, and F1-score.\"}]","A feature selection and scoring scheme for dimensionality reduction in a machine learning task | PDF",1785820716,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-feature-selection-and-scoring-scheme-for-dimensionality-reduction-in-a-machine-learning-task","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-feature-selection-and-scoring-scheme-for-dimensionality-reduction-in-a-machine-learning-task/124149/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in machine learning tasks?","Question",{"text":75,"@type":76},"The paper addresses challenges of handling high-dimensional datasets with many features, including noisy, redundant, and uninformative variables that can degrade learning performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method select important features?",{"text":80,"@type":76},"It uses a statistical lift measure to score and select an important feature subset, designed as a generic approach for binary classification datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"Which baselines and evaluation metrics are used to compare techniques?",{"text":84,"@type":76},"The proposed technique is compared with Chi-Square, Pearson Correlation, and Information Gain, and evaluated across five machine learning models using accuracy, precision, recall, and F1-score.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]