[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123210-en":3,"doc-seo-123210-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123210,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Optimized colon cancer classification via feature selection and machine learning - Research article","Optimized colon cancer classification uses gene-expression microarray data to address the high-dimensionality and limited-sample challenge that degrades prediction quality. A filtering approach (FA) and a gene classifier (GC) are combined to improve gene selection while enhancing classification accuracy. Experiments on 62 samples integrate statistical measures with multiple machine-learning classifiers, reaching about 96% and 97% accuracy. Results show that FA removes noise and redundancy, enabling accurate predictions from a minimal gene subset. The study supports robust feature selection for improved cancer diagnostics and motivates adaptive strategies for future cancer genomics research, including personalized medicine applications.","Optimized colon cancer classification via feature selection and  \nmachine learning  \nSara Haddou Bouazza, Jihad Haddou Bouazza  \nResearch Laboratory of the Moroccan School of Engineering Sciences, LAMIGEP, EMSI, Marrakech, Morocco  \nArticle history:  \nReceived Sep 10, 2024 Revised Oct 25, 2024 Accepted Nov 19, 2024  \nKeywords:  \nArtificial intelligence Cancer classification Computer science Feature selection Machine learning  \nCorresponding Author:  \nThe increasing dimensionality of gene expression data poses significant challenges in cancer classification, particularly in colon cancer. This study presents a novel filtering approach (FA) and a gene classifier (GC) to enhance gene selection and classification accuracy. Utilizing a dataset of 62 samples, our methods integrate statistical measures and machine learning classifiers, achieving classification accuracies of 96% and 97%, respectively. The FA effectively filters out noise and redundancy, allowing for accurate predictions with a minimal subset of genes, while the GC leverages multiple classifiers for optimal performance. These findings underscore the importance of robust feature selection in improving cancer diagnostics and suggest potential applications in personalized medicine. By addressing the limitations of existing methodologies, our work lays the groundwork for future research in cancer genomics, emphasizing the need for adaptive strategies to handle complex datasets.  \nThis is an open access article under the CC BY-SA license.  \nSara Haddou Bouazza  \nResearch Laboratory of the Moroccan School of Engineering Sciences, LAMIGEP, EMSI Marrakech, Morocco  \nEmail: [sara.hb.sara@gmail.com](sara.hb.sara@gmail.com)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nDeoxyribonucleic acid (DNA) microarray technology has transformed cancer research, enabling simultaneous analysis of thousands of genes, and offering valuable insights into gene interactions crucial for early detection, diagnosis, and prognosis [1], [2] . Despite this, significant challenges remain due to the imbalance between the large number of genes and the limited sample size in such datasets. Many genes are irrelevant to cancer progression or highly interdependent, complicating analyses and potentially leading to inaccurate predictions if the entire gene set is used indiscriminately [3], [4] .  \nFeature selection is pivotal in overcoming these challenges by reducing dimensionality and excluding irrelevant or noisy genes, enhancing classification accuracy and model interpretability [5], [6] . Recent advancements in feature selection methods have been made. For instance, Hegazy et al. [7] demonstrated the efficacy of differential evolution techniques for colon cancer gene selection , while Ali and Saeed [8] developed a hybrid filter-genetic algorithm (GA) that improved classification across various cancer types.  \nOther studies, including those by Kourou et al. [9] and Hambali et al. [10], highlighted the importance of feature selection in cancer classification, noting the need for more refined techniques to address the complexity of microarray data. Chowdhary et al. [11] further identified ongoing issues in featureselection for high-dimensional datasets. Hybrid methods, such as those combining filter methods with algorithms like C5.0, have shown promise, as illustrated by Hamim et al. [12], who achieved significant improvements in breast cancer classification. Despite these advancements, current feature selection methods  \nrequire further refinement to balance classification accuracy, computational efficiency, and gene interpretability, especially in colon cancer classification. This study addresses these gaps by proposing an innovative approach integrating filter and wrapper methods to identify the most relevant genes, aiming to enhance both accuracy and efficiency in cancer diagnostics.  \nThe structure of this paper is organized as follows: section 2 discusses the materials and methods used, detailing the datase","cbCaiss3tehhETQU","https://ap.wps.com/l/cbCaiss3tehhETQU","pdf",572115,1,10,"English","en",105,"# Introduction\n## Feature selection motivation and related work\n# Method\n## Three-step gene selection approach (filter, wrapper, minimal subset)\n## Filter approach (SNR, Pearson correlation, ReliefF)\n# Results\n## Accuracy metrics and comparisons (state-of-the-art)\n# Discussion\n## Findings, limitations, and future work\n# Conclusion\n## Summary of contributions","[{\"question\":\"Why is feature selection important for colon cancer classification with gene expression data?\",\"answer\":\"Gene expression microarrays produce thousands of genes but only a limited number of samples, making models vulnerable to irrelevant, redundant, or noisy features. Feature selection reduces dimensionality, improves accuracy, and enhances interpretability by excluding unhelpful genes.\"},{\"question\":\"How does the proposed approach combine filtering and classification?\",\"answer\":\"The method first applies a filter approach (FA) using statistical measures such as SNR and Pearson correlation to remove noise and redundancy. It then applies a wrapper refinement and finally keeps a minimal subset of genes to balance accuracy with computational efficiency before classification.\"},{\"question\":\"What classification performance is reported for the dataset used in this study?\",\"answer\":\"Using a dataset of 62 samples, the integrated methods achieve classification accuracies of approximately 96% for the FA-based filtering pipeline and about 97% for the gene classifier strategy that leverages multiple classifiers for optimal performance.\"}]","Optimized colon cancer classification via feature selection and machine learning - Research article | PDF",1785815229,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"optimized-colon-cancer-classification-via-feature-selection-and-machine-learning-research-article","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/optimized-colon-cancer-classification-via-feature-selection-and-machine-learning-research-article/123210/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is feature selection important for colon cancer classification with gene expression data?","Question",{"text":76,"@type":77},"Gene expression microarrays produce thousands of genes but only a limited number of samples, making models vulnerable to irrelevant, redundant, or noisy features. Feature selection reduces dimensionality, improves accuracy, and enhances interpretability by excluding unhelpful genes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed approach combine filtering and classification?",{"text":81,"@type":77},"The method first applies a filter approach (FA) using statistical measures such as SNR and Pearson correlation to remove noise and redundancy. It then applies a wrapper refinement and finally keeps a minimal subset of genes to balance accuracy with computational efficiency before classification.",{"name":83,"@type":74,"acceptedAnswer":84},"What classification performance is reported for the dataset used in this study?",{"text":85,"@type":77},"Using a dataset of 62 samples, the integrated methods achieve classification accuracies of approximately 96% for the FA-based filtering pipeline and about 97% for the gene classifier strategy that leverages multiple classifiers for optimal performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":21,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]