[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160237-en":3,"doc-seo-160237-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160237,962084925636,"Sophia Brooks","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Using PCA and Factor Analysis for Dimensionality Reduction of Bio-informatics Data","Large genomics datasets generated by advances in sequencing require proper analysis to become useful. Bioinformatics data typically contains hundreds of high-dimensional attributes, degrading machine learning performance for classification and prediction. This study applies Principal Component Analysis and Factor Analysis to reduce dimensionality and improve downstream analysis. Experiments use a leukemia dataset, aiming to reduce the number of attributes while retaining relevant information for effective statistical testing and further machine learning tasks.","Using PCA and Factor Analysis for Dimensionality Reduction of Bio-informatics Data  \nM. Usman Ali  \nDepartment of Computer Science COMSATS Institute of Information Technology Sahiwal, Pakistan  \nShahzad Ahmed  \nDepartment of Computer Science COMSATS Institute of Information Technology Sahiwal, Pakistan  \nJaved Ferzund  \nDepartment of Computer Science COMSATS Institute of Information Technology Sahiwal, Pakistan  \nAtif Mehmood  \nRiphah Institute of Computing and Applied Sciences (RICAS) Riphah International University Lahore, Pakistan  \nAbbas Rehman  \nDepartment of Computer Science COMSATS Institute of Information Technology Sahiwal, Pakistan  \nAbstract—Large volume of Genomics data is produced on daily basis due to the advancement in sequencing technology. This data is of no value if it is not properly analysed. Different kinds of analytics are required to extract useful information from this raw data. Classification, Prediction, Clustering and Pattern Extraction are useful techniques of data mining. These techniques require appropriate selection of attributes of data forgetting accurate results. However, Bioinformatics data is high dimensional, usually having hundreds of attributes. Such large a number of attributes affect the performance of machine learning algorithms used for classification/prediction. So, dimensionality reduction techniques are required to reduce the number of attributes that can be further used for analysis. In this paper, Principal Component Analysis and Factor Analysis are used for dimensionality reduction of Bioinformatics data. These techniques were applied on Leukaemia data set and the number of attributes was reduced from to.  \nKeywords—Bioinformatics; Statistics; Microarray; Leukaemia; Feature Selection; Statistical tests; PCA; Factor Analysis; R tool  \nI. INTRODUCTION  \nBioinformatics experiments are based on Genome, DNA, RNA and Chromosomes. Genomics plays an imperative role in this field. Huge amount of data has been produced in Genomics with a substantial portion produced in Functional Genomics (in the form of protein-protein association), Structural Genomics (in the form of 3-D structure). By using NGS (Next Generation Sequencing) technique, a lot of work has been done in the field of Microarray. This technique is helpful to identify human diseases. NGS is sequencing technique which is used to detect sequences of proteomics for next generation.  \nGenetic Diseases are caused by Genetic disorders that are more complex because of multiple genes interaction. These disorders are breast cancer, colon cancer, skin cancer, autism, progeria, and haemophilia. These are caused by mutation in genes or sometimes inherit from parents. Leukaemia is a cancer of blood cells that occurs due to genome abnormality. Microarray includes genes expression data that is present at large scale.  \nBioinformatics data needs to be store in an efficient manner and include a lot of Attributes (Variables) . The major problem is that, most of tools crash when large data stored in it.  \nStatistics plays superlative role in the field of Bioinformatics, Mathematics and Computer Science. It is used to extract, organise, analyse and visualise large amount of data.  \n. For this purpose, a lot of tools like Excel, Weka, Matlab and R are available. Many Statistical tests are used for the extraction of relevant information. These are t-test, chi-squared test (χ2-test), ANOVA (Analysis of Variance), Kruskal-Wallis, Friedman and PCA (Principle Component Analysis) tests [1] Statistical t-test is used to check the difference between sample Mean and hypothesised value. ANOVA is parametric (distribution) test used to check the difference of dependent variables with levels of independent variables. Kruskal-Wallisis non-parametric (distribution free) test in which assumptions are not including unlike ANOVA. Friedman test is used when there is one distributed dependent variable and one independent variable with two or many levels. It is used to","cbCaienLlJWjvcAm","https://ap.wps.com/l/cbCaienLlJWjvcAm","pdf",1035448,1,12,"English","en",105,"# Introduction\n## Bioinformatics and high-dimensional genomics data\n## Statistical tests and analysis tools\n## R for bioinformatics analytics\n# Related Work\n# Experimental Setup\n# Results and Discussion\n# Conclusion","[{\"question\":\"Why is dimensionality reduction necessary for bioinformatics data?\",\"answer\":\"Bioinformatics datasets are high-dimensional, often containing hundreds of attributes. This large number of variables can reduce the performance of machine learning algorithms used for classification and prediction, so dimensionality reduction is needed.\"},{\"question\":\"What techniques does the paper use for dimensionality reduction?\",\"answer\":\"The paper uses Principal Component Analysis (PCA) and Factor Analysis to reduce the number of attributes in bioinformatics data so that more effective analysis can be performed.\"},{\"question\":\"Which dataset is used to apply PCA and Factor Analysis?\",\"answer\":\"The techniques are applied on a leukemia dataset, where the goal is to reduce the number of attributes and extract more relevant information for downstream tasks.\"}]","Using PCA and Factor Analysis for Dimensionality Reduction of Bio-informatics Data | PDF",1788052650,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"using-pca-and-factor-analysis-for-dimensionality-reduction-of-bio-informatics-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/using-pca-and-factor-analysis-for-dimensionality-reduction-of-bio-informatics-data/160237/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is dimensionality reduction necessary for bioinformatics data?","Question",{"text":75,"@type":76},"Bioinformatics datasets are high-dimensional, often containing hundreds of attributes. This large number of variables can reduce the performance of machine learning algorithms used for classification and prediction, so dimensionality reduction is needed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What techniques does the paper use for dimensionality reduction?",{"text":80,"@type":76},"The paper uses Principal Component Analysis (PCA) and Factor Analysis to reduce the number of attributes in bioinformatics data so that more effective analysis can be performed.",{"name":82,"@type":73,"acceptedAnswer":83},"Which dataset is used to apply PCA and Factor Analysis?",{"text":84,"@type":76},"The techniques are applied on a leukemia dataset, where the goal is to reduce the number of attributes and extract more relevant information for downstream tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]