[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120420-en":3,"doc-seo-120420-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120420,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","The Classification of Cancer Subtype Based on Machine Learning - Bachelor’s Thesis 2025","Effective cancer subtype classification supports more personalized treatment planning and improved patient outcomes, yet traditional approaches may not capture the full complexity and diversity of cancer types. This study applies machine learning to TCGA gene expression profiles to improve subtype identification, using three base classifiers (RF, SVM, KNN) and a stacked ensemble with logistic regression as meta-classifier. Results evaluate multi-class performance with accuracy and macro/weighted precision, recall, and F1. The stacked ensemble achieves the best overall metrics and offers a scalable bioinformatics solution.","The Classification of Cancer Subtype Based on Machine Learning  \nLappeenranta–Lahti University of Technology LUT  \nBachelor's thesis  \nBachelor’s programme in Software and Systems Engineering  \n2025  \nZhihao Liu  \nSupervisors: S.E. Yao Dong  \nDr. Sonja Hyrynsalmi  \nABSTRACT  \nLappeenranta–Lahti University of Technology LUTLUT School of Engineering Sciences  \nSoftware and Systems Engineering  \nIn co-operation with partner university/universities: Hebei University of Technology  \nZhihao Liu  \nThe Classification of Cancer Subtype Based on Machine Learning  \nBachelor’s thesis 2025  \n39 pages, 10 figures, 1 table  \nSupervisors: S.E. Yao Dong, Dr. Sonja Hyrynsalmi  \nKeywords: Cancer, Classification, Machine Learning, Ensemble Model, Logical Regression  \nEffective classification of cancer subtypes plays a crucial role in designing personalized treatment plans and improving patient outcomes. Traditional classification methods may not reflect the complexity and diversity of cancer types. This study utilizes gene expression profiles from TCGA to investigate how machine learning techniques can be improved to identify cancer subtypes.  \nThe study used three machine learning algorithms (RF, SVM, KNN), as well as a stacked ensemble approach using logistic regression as a meta-classifier. The study processed gene expression data from 10,446 samples of 33 different cancer types from TCGA, and their type labels. Model performance was evaluated using various metrics including overall accuracy, macro-mean of precision, macro-mean of recall, macro-mean ofF1 score, weighted mean of precision, weighted mean of recall, and weighted mean ofF1 score.  \nThe results show that the stacked ensemble model outperforms the individual classifiers, achieving an accuracy of 0.7836, a macroscopic precision of 0.7257, a macroscopic recall of 0.6746, and a macroscopic F1 score of 0.6860, with values of 0.7713 for the weighted precision, 0.7763 for the recall, and 0.7637 for the F1 score, respectively. Ensemble methods successfully utilize different classification boundaries for RF, SVM, and KNN to ensure more reliable classification for less common subtypes. These results emphasize the benefits of integrated learning methods in bioinformatics and present a scalable solution for cancer subtype prediction. Future work could enhance this framework by incorporating multi-omics datasets and validating results with external clinical resources.  \nACKNOWLEDGEMENTS  \nI am especially grateful to S.E. Yao Dong and Dr. Sonja Hyrynsalmi for their guidance on my research direction and thesis writing.  \nI would like to show my sincere gratitude to TCGA program for providing reliable data resources for this research. TCGA has made great contributions to the advancement of human cancer genomics research, providing systematic and detailed data support to researchers around the globe, which has made it possible to carry out this study.  \nMeanwhile, I would like to thank all the researchers in cancer classification and its machine learning methods. It is their solid and cutting-edge work that provides a solid theoretical foundation and practical guidance for my research.  \nABBREVIATIONS  \nWHO World Health Organization  \nRF Random Forest  \nKNN K-Nearest Neighbour  \nSVM Support Vector Machine  \nTCGA The Cancer Genome Atlas  \nML Machine Learning  \nCNN Convolutional Neural Networks  \nTable of contents  \nAbstract  \nAcknowledgements  \nAbbreviations  \nTable of Contents  \n1 Introduction .................................................................................................................... 1  \n1.1 Background ............................................................................................................. 1  \n1.2 Study Purpose.......................................................................................................... 2  \n1.3 Thesis Structure....................................................................................................... 2  \n2 Relate","cbCaiopOPkvnkwPF","https://ap.wps.com/l/cbCaiopOPkvnkwPF","pdf",730443,1,39,"English","en",105,"# Abstract\n# Acknowledgements\n# Abbreviations\n# 1 Introduction\n## 1.1 Background\n## 1.2 Study Purpose\n## 1.3 Thesis Structure\n# 2 Related Research\n## 2.1 Cancer Subtype\n## 2.2 Machine Learning\n## 2.3 The Application of Machine Learning for Cancer Subtype Classification\n## 2.4 Stacking Generalization\n## 2.5 Data Processing in Cancer Subtype Classification\n## 2.6 Evaluation Metrics for Subtype\n# 3 Research Method\n## 3.1 Experiment Environment\n## 3.2 Overview of Research Method\n## 3.3 Data Collection & Preprocessing\n## 3.4 Model Selection\n## 3.5 Model Training","[{\"question\":\"What dataset and labeling approach are used to train the cancer subtype models?\",\"answer\":\"The study uses gene expression data from TCGA with 10,446 samples covering 33 cancer types. Each sample uses its cancer type label for supervised learning.\"},{\"question\":\"Which machine learning models are compared in the study?\",\"answer\":\"Three base classifiers are used: Random Forest (RF), Support Vector Machine (SVM), and K-Nearest Neighbour (KNN). A stacked ensemble combines them using logistic regression as the meta-classifier.\"},{\"question\":\"How does the stacked ensemble perform compared with individual classifiers?\",\"answer\":\"The stacked ensemble outperforms the individual models, reaching an accuracy of 0.7836 and macro-level precision, recall, and F1 scores of 0.7257, 0.6746, and 0.6860, along with strong weighted metrics.\"}]","The Classification of Cancer Subtype Based on Machine Learning - Bachelor’s Thesis 2025 | PDF",1785729955,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-classification-of-cancer-subtype-based-on-machine-learning-bachelors-thesis-2025","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-classification-of-cancer-subtype-based-on-machine-learning-bachelors-thesis-2025/120420/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What dataset and labeling approach are used to train the cancer subtype models?","Question",{"text":75,"@type":76},"The study uses gene expression data from TCGA with 10,446 samples covering 33 cancer types. Each sample uses its cancer type label for supervised learning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are compared in the study?",{"text":80,"@type":76},"Three base classifiers are used: Random Forest (RF), Support Vector Machine (SVM), and K-Nearest Neighbour (KNN). A stacked ensemble combines them using logistic regression as the meta-classifier.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the stacked ensemble perform compared with individual classifiers?",{"text":84,"@type":76},"The stacked ensemble outperforms the individual models, reaching an accuracy of 0.7836 and macro-level precision, recall, and F1 scores of 0.7257, 0.6746, and 0.6860, along with strong weighted metrics.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]