[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124122-en":3,"doc-seo-124122-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124122,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Machine learning methods for biological hypothesis generation - facilitating new discoveries at lower costs - dissertation","Machine learning methods for biological hypothesis generation focus on overcoming two bottlenecks in modern biology: expensive experimental measurements and data-hungry algorithms that often lack large public datasets. The thesis leverages computational simulation with generative models to synthesize additional in silico datapoints from heterogeneous high-quality sources. Three methods are developed across genomic time-series, dense chromatin contact maps, and gene network construction. Each model is tied back to practical biology workflows, supporting both expert-assisted analysis and novel hypothesis prioritization for later lab testing.","Machine learning methods for biological hypothesis generation, facilitating new discoveries at lower costs  \nAdelaide Woods Chambers Woicik  \nA dissertation  \nsubmitted in partial fulfillment of the  \nrequirements for the degree of  \nDoctor of Philosophy  \nUniversity of Washington  \n2024  \nReading Committee:  \nSheng Wang, Chair  \nBill Noble  \nLinda G. Shapiro  \nProgram Authorized to Offer Degree:  \nComputer Science and Engineering  \n©Copyright 2024  \nAdelaide Woods Chambers Woicik  \nUniversity of Washington  \nAbstract  \nMachine learning methods for biological hypothesis generation, facilitating new discoveries at lower costs  \nAdelaide Woods Chambers Woicik  \nChair of the Supervisory Committee:  \nSheng Wang  \nComputer Science and Engineering  \nMachine learning methods for biological data have become increasingly popular in recent years, acknowledging the transformative applications, complex patterns, and latent variation underlying biological systems. Importantly, many biological measurements are very expensive to produce experimentally. This poses challenges for biological discovery, limiting the number of experiments that can practically be conducted, and for data-hungry machine learning methods, which may require massive datasets that are not publicly available. One approach to these challenges is computational simulation with generative machine learning models, leveraging available high-quality data from heterogeneous sources to synthesize additional datapoints for subsequent analyses, which can help propose novel and prioritize existing biological hypotheses that can subsequently be tested in an experimental lab. In this thesis, I present three methods for highquality in silico data generation across three biological domains: genomic time series extrapolation with Sagittarius, high-resolution dense chromatin contact map generation with Capricorn, and approximatelyautomatically-curated gene network generation using augmented network integration with Gemini. These diverse applications focus on high-cost experimental data, highlighting the immense value of computational datapoint simulation, and heterogeneous biological measurements, requiring methods that account for the diverse inputs and leverage all sources of information to improve the generation process. Finally, I connect each model back to its practical applications in biology, ranging from assisting biological experts in their current work to novel hypothesis generation.  \n4  \nAcknowledgements  \nI’d like to thank the many people who helped make the work in this thesis possible. First and foremost, I’d like to thank my advisor, Professor Sheng Wang, for his guidance during the PhD. I’d also like to thank my other committee members Professors William Noble, Linda Shapiro, Simon Du, and Daniela Witten for their helpful feedback and support.  \nI’d also like to thank the members of the Wang lab for the useful collaborations and discussions throughout my tenure at the University of Washington. In particular, I’d like to thank Hanwen Xu, Tangqi Fang, Yifeng Liu, Mingxin Zhang, Zixuan Liu, Tong Chen, and Guang Yang. Furthermore, I’d like to acknowledge the members of the working computational Hi-C modeling group, including Anupama Jha, Xiao Wang, Borislav Hristov, Gang Li, and Shengqi Hang, for teaching me so much about underlying cellular biology and for all of their advice in the past two years. Finally, I’d like to thank Doctors Jay Pal and Shin Lin for their support in the past year, and for teaching me about various opportunities for generative machine learning in clinical settings, and particularly in cardiology.  \nDEDICATION  \nTo my family: my husband, Matt Woicik; my dad, Craig Chambers; and my sister, Caitlin McWilliams.  \nAnd for my mom, Sylvia Chambers.  \nAll my love, always.  \n8  \nContents  \n1 Introduction 19  \n1.1 Biological datasets: challenges and opportunities ....................... 21  \n1.2 In silico generation of high-cost biological data ................","cbCaihgAJqyFqnTT","https://ap.wps.com/l/cbCaihgAJqyFqnTT","pdf",57341894,1,186,"English","en",105,"# Contents\n## Introduction\n## Background\n## Sagittarius: Biological time series extrapolation","[{\"question\":\"为什么本研究关注“以较低成本提出新的生物假设”？\",\"answer\":\"因为许多生物实验测量成本很高且数量有限，而数据驱动的机器学习往往需要大规模数据。研究用计算生成来补足高成本实验的不足。\"},{\"question\":\"论文提出了哪三类用于生物数据生成的方法与领域？\",\"answer\":\"分别面向基因组时间序列外推、致密染色质接触图（Hi-C）生成，以及用于构建/策划基因网络的增强网络集成流程。\"},{\"question\":\"这些方法如何帮助后续实验室验证与发现？\",\"answer\":\"生成的额外 in silico 数据可用于提出新的生物假设并优先排序，再由实验室进行检验，从而加速发现流程。\"}]","Machine learning methods for biological hypothesis generation - facilitating new discoveries at lower costs - dissertation | PDF",1785820565,469,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-methods-for-biological-hypothesis-generation-facilitating-new-discoveries-at-lower-costs-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-methods-for-biological-hypothesis-generation-facilitating-new-discoveries-at-lower-costs-dissertation/124122/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么本研究关注“以较低成本提出新的生物假设”？","Question",{"text":75,"@type":76},"因为许多生物实验测量成本很高且数量有限，而数据驱动的机器学习往往需要大规模数据。研究用计算生成来补足高成本实验的不足。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"论文提出了哪三类用于生物数据生成的方法与领域？",{"text":80,"@type":76},"分别面向基因组时间序列外推、致密染色质接触图（Hi-C）生成，以及用于构建/策划基因网络的增强网络集成流程。",{"name":82,"@type":73,"acceptedAnswer":83},"这些方法如何帮助后续实验室验证与发现？",{"text":84,"@type":76},"生成的额外 in silico 数据可用于提出新的生物假设并优先排序，再由实验室进行检验，从而加速发现流程。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]