[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121290-en":3,"doc-seo-121290-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},121290,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Advancing Cancer Classifcation through Machine Learning Analysis of RNA-Seq Gene Expression Data","This study advances cancer classification using the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, leveraging gene expression profiles across multiple tumor samples. The work addresses high-dimensional learning challenges, including the Hughes Effect and the Curse of Dimensionality, via feature selection and machine learning. Tree-based modeling with Random Forest refines the problem to seventy highly relevant genes, while PCA and Kernel PCA enable visualization of non-linear gene-expression patterns. Network analysis further examines gene interaction communities. Evaluation compares SVM, Logistic Regression, and KNN, emphasizing precision and low test error rates, supporting systematic analysis for future personalized medicine.","University of Central Florida  \nSTARS  \nData Science and Data Mining  \nSpring 2024  \nAdvancing Cancer Classifcation through Machine Learning Analysis of RNA-Seq Gene Expression Data  \nEmil Agbemade  \nUniversity of Central Florida, [emil.agbemade@ucf.edu](emil.agbemade@ucf.edu)  \nAmina Issoufou Anaroua  \nUniversity of Central Florida, [amina.issoufouanaroua@ucf.edu](amina.issoufouanaroua@ucf.edu)  \nDimitri Bamba  \nUniversity of Central Florida, [di714545@ucf.edu](di714545@ucf.edu)  \n Part of the Data Science Commons  \nFind similar works at: [https://stars.library.ucf.edu/data-science-mining](https://stars.library.ucf.edu/data-science-mining)  \nUniversity of Central Florida Libraries [http://library.ucf.edu](http://library.ucf.edu)  \nThis Article is brought to you for free and open access by STARS. It has been accepted for inclusion in Data Science and Data Mining by an authorized administrator of STARS. For more information, please [contact STARS@ucf.edu](contact STARS@ucf.edu).  \nSTARS Citation  \nAgbemade, Emil; Anaroua, Amina Issoufou; and Bamba, Dimitri, \"Advancing Cancer Classifcation through Machine Learning Analysis of RNA-Seq Gene Expression Data\" (2024) . Data Science and Data Mining. 16.  \n[https://stars.library.ucf.edu/data-science-mining/16](https://stars.library.ucf.edu/data-science-mining/16)  \nAdvancing Cancer Classifcation through Machine Learning Analysis of RNA-Seq Gene Expression  \nData  \n1st Emil Agbemade  \nDept. of Statistics and Data Science  \nUniversity of Central Florida  \nOrlando, USA  \n[emil.agbemade@ucf.edu](emil.agbemade@ucf.edu)  \nAbstract—This study delves into the classifcation of various cancer types using the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, which encompasses a rich collection of gene expression data across multiple tumor samples. To improve cancer diagnosis and treatment, our methodology confronts the challenges inherent in high-dimensional datasets, such as the Hughes Effect and the Curse of Dimensionality, through innovative feature selection methods and machine learning approaches. A key component of our strategy includes the use of tree-based algorithms, particularly Random Forest, to refne the dataset to seventy genes of utmost relevance for tumor classifcation, and the application of PCA and Kernel PCA for dimensional reduction, enabling the visualization of non-linear patterns in gene expression data. The research further investigates the gene interaction network through network analysis, employing modularity metrics to understand signifcant community structures linked to biological processes in cancer. Our model evaluation assesses various machine learning models, highlighting the precision and low-test error rates of SVM, Logistic Regression, and KNN, suggesting their effectiveness in exploiting the dataset’s inherent separability. The study’s comprehensive approach not only provides a systematic framework for analyzing gene expression data but also paves the way for advanced research into the genetic mechanisms of cancer, with implications for personalized medicine and treatment strategies.  \nIndex Terms—Cancer classifcation, gene expression data, RNA-Seq, machine learning, feature selection, dimensionality reduction, network analysis.  \nI. BACKGROUND AND MOTIVATION  \nMicroarrays and RNA-Seq are two examples of the cuttingedge biotech tools that have allowed researchers to capture gene expression data from tissues. In order to improve diagnosis and treatment optimization, these methodologies allow for the separation between healthy and sick states, as well as between different types and subtypes of cancer [1], [2],[25] . But there are obstacles to cancer categorization, like the Hughes Effect and the Curse of Dimensionality [5], [6] . This problem occurs when there are more genes in DNA sequences than there are samples, which causes classifcation algorithms to be less effective because of irrelevant genes [7], [8] .  \nTo tackle this, a lot of res","cbCaipmT3ReUc2zQ","https://ap.wps.com/l/cbCaipmT3ReUc2zQ","pdf",899522,1,"English","en",105,"# Background and Motivation\n## Feature Selection Challenges\n## Prior Approaches and Gene Marker Selection\n## Public RNA-Seq Datasets and Related Results\n# Methodology Overview\n## Data Source and RNA-Seq Dataset\n## Feature Selection with Tree-Based Models\n## Dimensionality Reduction via PCA and Kernel PCA\n## Gene Interaction Network Analysis\n# Model Evaluation\n## Compared Classifiers and Metrics\n## Performance Highlights\n# Conclusion and Implications","[{\"question\":\"Which dataset and data type are used for the cancer classification task?\",\"answer\":\"The study uses the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, containing gene expression data across multiple tumor samples.\"},{\"question\":\"How does the study handle high-dimensionality problems like the Hughes Effect and the Curse of Dimensionality?\",\"answer\":\"It applies feature selection to reduce irrelevant genes, then uses tree-based methods (Random Forest) to narrow the set to seventy highly relevant genes and employs dimensionality reduction (PCA and Kernel PCA).\"},{\"question\":\"Which machine learning models are evaluated and what performance aspects are highlighted?\",\"answer\":\"The research evaluates SVM, Logistic Regression, and KNN, emphasizing precision and low test error rates as indicators of their effectiveness.\"}]","Advancing Cancer Classifcation through Machine Learning Analysis of RNA-Seq Gene Expression Data | PDF",1785734926,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"advancing-cancer-classifcation-through-machine-learning-analysis-of-rna-seq-gene-expression-data","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/advancing-cancer-classifcation-through-machine-learning-analysis-of-rna-seq-gene-expression-data/121290/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Which dataset and data type are used for the cancer classification task?","Question",{"text":74,"@type":75},"The study uses the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, containing gene expression data across multiple tumor samples.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the study handle high-dimensionality problems like the Hughes Effect and the Curse of Dimensionality?",{"text":79,"@type":75},"It applies feature selection to reduce irrelevant genes, then uses tree-based methods (Random Forest) to narrow the set to seventy highly relevant genes and employs dimensionality reduction (PCA and Kernel PCA).",{"name":81,"@type":72,"acceptedAnswer":82},"Which machine learning models are evaluated and what performance aspects are highlighted?",{"text":83,"@type":75},"The research evaluates SVM, Logistic Regression, and KNN, emphasizing precision and low test error rates as indicators of their effectiveness.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]