[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117906-en":3,"doc-seo-117906-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117906,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Evaluation of Machine Learning Methods to Decode Transcriptional Regulation","High-throughput technologies enable large-scale genomic measurements and support epigenetic profiling of transcriptional regulation. This thesis investigates how sequence fragments from an ATAC-STARR-seq experiment can drive transcription by predicting log2FoldChange using over five million fragments from a salmon genome, with training data selected from the top 10% basemean. Three supervised regression approaches are evaluated, including XGBoost, Random Forest, and Linear SVR, emphasizing feature engineering and feature importance for motif search, yielding model-derived sequence features linked to known transcription factors.","Master’s Thesis 2023 30 ECTS  \nFaculty of Science and Technology  \nEvaluation of Machine Learning Methods to Decode Transcriptional Regulation  \nHarini Jeyakumar Data Science  \nAcknowledgement  \nFirstly, I want to thank all my supervisors for their guidance and help during the work of this thesis. I want to thank my main supervisor Dean, Torgeir Rhoden Hvidsten for introducing me to this topic. I also want to thank my co-supervisors Prof. Simen Rød Sandve and Researcher Lars Grønvold for sharing their knowledge in biology and for giving me advice. In addition, I would like to thank my co-supervisor Prof. Kristian Hovde Liland for technical help whenever I needed it.  \nFurthermore, I would like to thank my family and close friends for their support, especially my roommates for interesting ideas throughout this semester.  \nHarini Jeyakumar, May 2023  \nii  \nAbstract  \nWith large biological measurements made possible by the development of high-throughput technology, it allows for the study of genomic data. In transcriptional regulation the cell controls the translation of DNA to RNA, and thereby controls which genes to express. Exploring the regulatory mechanisms underlying the genes that are controlled in a cell is part of epigenetic profiling. Finding accessible DNA regions that control transcriptional regulation can be done using this method. It appears that machine learning models have not previously been tested on ATAC-STARR-seq data with varying fragment length from the salmon genome. Using the results from an ATAC-STARR-seq experiment on a salmon genome, we want to explore the extent to which it is possible to predict from sequence fragments.  \nIn this thesis, we attempt to predict to what degree sequence fragments can drive transcription using the ATAC-STARR-seq results of over five million sequence fragments. Fragments within the top 10% basemean were selected for the model’s training, validation, and testing in order to improve the performance of the machine learning models. Techniques such as feature engineering and feature importance were crucial for training the model and extracting relevant features for motif search. We evaluated three classical machine learning methods for regression analysis predicting log2FoldChange values. The ensemble methods XGBoost Regressor and Random Forest Regressor, and Linear Support Vector Regression were applied as machine learning algorithms. XGBoost Regressor and Random Forest Regressor are both powerful algorithms known to have been used successfully on sequential data. Linear Support Vector Regression’s ability to handle a large number of samples compared to Support Vector Regression, was one of the reasons this algorithm was chosen.  \nFurthermore, XGBoost Regressor and Random Forest Regressor performed remarkably similar with potential for improvement. The Linear Support Vector Regression model had more trouble capturing the complexity of the data. Despite the results, it was possible to extract sequence features from the trained machine learning models and associate them with known transcription factors. This study’s findings indicate that there is still a need for improvements in the performance of the models. While time limit and time-consuming algorithms have been a challenge, possibilities for tuning with not yet tested parameters and other methods to improve the model remains for further research.  \nContents  \nPreface i  \nAbstract iii  \nTable of Contents iv  \nList of Figures vi  \n1 Introduction 1  \n1.1 Motivation ....................................... 1  \n1.2 Objectives ........................................ 2  \n1.3 Structure ........................................ 2  \n2 Theory 3  \n2.1 DNA ........................................... 3  \n2.1.1 Transcription and Translation ......................... 4  \n2.1.2 Gene Regulation ................................ 4  \n2.2 ATAC-STARR-seq ................................... 6  \n2.3 Machine Learning ..................................","cbCairf38IWGhO3P","https://ap.wps.com/l/cbCairf38IWGhO3P","pdf",1826576,1,46,"English","en",105,"# 1 Introduction\n## 1.1 Motivation\n## 1.2 Objectives\n## 1.3 Structure\n# 2 Theory\n## 2.1 DNA\n## 2.2 ATAC-STARR-seq\n## 2.3 Machine Learning\n# 3 Method\n## 3.1 Dataset\n## 3.2 Preprocessing\n## 3.3 XGBoost Regressor Model\n## 3.4 Random Forest Regressor Model\n## 3.5 Linear SVR\n# 4 Results\n## 4.1 XGBoost Regressor Default VS Tuned\n## 4.2 Random Forest Regressor: Default VS Tuned\n## 4.3 Linear SVR: Default VS Tuned\n## 4.4 Motif Search Discovery\n# 5 Discussion\n## 5.1 Data Quality\n## 5.2 Model Performance\n## 5.3 Future Work\n# 6 Conclusion\n## A Table of Python libraries","[{\"question\":\"What biological problem does the thesis address?\",\"answer\":\"It studies transcriptional regulation and how genomic sequence fragments can be used to predict which genes are expressed, using ATAC-STARR-seq results for epigenetic profiling.\"},{\"question\":\"Which machine learning methods are evaluated and what are they predicting?\",\"answer\":\"Three supervised regression methods are tested to predict log2FoldChange values: XGBoost Regressor, Random Forest Regressor, and Linear Support Vector Regression.\"},{\"question\":\"How were training samples selected before model training?\",\"answer\":\"Fragments were filtered by basemean, selecting those within the top 10% basemean for training, validation, and testing to improve model performance.\"}]","Evaluation of Machine Learning Methods to Decode Transcriptional Regulation | PDF",1785680308,116,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"evaluation-of-machine-learning-methods-to-decode-transcriptional-regulation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/evaluation-of-machine-learning-methods-to-decode-transcriptional-regulation/117906/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What biological problem does the thesis address?","Question",{"text":76,"@type":77},"It studies transcriptional regulation and how genomic sequence fragments can be used to predict which genes are expressed, using ATAC-STARR-seq results for epigenetic profiling.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which machine learning methods are evaluated and what are they predicting?",{"text":81,"@type":77},"Three supervised regression methods are tested to predict log2FoldChange values: XGBoost Regressor, Random Forest Regressor, and Linear Support Vector Regression.",{"name":83,"@type":74,"acceptedAnswer":84},"How were training samples selected before model training?",{"text":85,"@type":77},"Fragments were filtered by basemean, selecting those within the top 10% basemean for training, validation, and testing to improve model performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]