[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126732-en":3,"doc-seo-126732-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126732,962084925782,"Ava Thompson","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Statistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems - Thesis","Biomedical research now faces rapidly expanding, heterogeneous datasets, creating both opportunity and risk for scientific discovery. This dissertation develops interpretable statistical machine learning methods for extracting reliable hypotheses from biomedical data through in-context method development, responsible data science practice, and dissemination of open-source software and data. Grounded in the Predictability, Computability, and Stability (PCS) framework, the work improves the trustworthiness of conclusions by promoting predictability checks, computability-aware design, and stability requirements for reproducibility and interpretability, with transparent documentation across the pipeline.","UC Berkeley  \nUC Berkeley Electronic Theses and Dissertations  \nTitle  \nStatistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems  \nPermalink  \n[https://escholarship.org/uc/item/42p9f9c8](https://escholarship.org/uc/item/42p9f9c8)  \nAuthor  \nTang, Tiffany  \nPublication Date  \n2023  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nStatistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems  \nby  \nTiffany Tang  \nA dissertation submitted in partial satisfaction of the requirements for the degree of  \nDoctor of Philosophy  \nin  \nStatistics  \nin the  \nGraduate Division  \nof the  \nUniversity of California, Berkeley  \nCommittee in charge:  \nProfessor Bin Yu, Chair  \nProfessor Haiyan Huang  \nAssistant Professor James B. Brown  \nSummer 2023  \nStatistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems  \nCopyright 2023  \nby  \nTiffany Tang  \n1  \nAbstract  \nStatistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems  \nby  \nTiffany Tang  \nDoctor of Philosophy in Statistics  \nUniversity of California, Berkeley  \nProfessor Bin Yu, Chair  \nGiven the ever-growing volume and variety of biomedical data, principled analyses of these rich datasets offer an exciting opportunity to accelerate the scientific discovery process. Here, we advance our goal of extracting reliable scientific hypotheses from such data through (I) the in-context development of interpretable statistical machine learning methods,(II) the demonstration of responsible data science in practice, and (III) the dissemination of opensource software and data for reliable data science.  \nThroughout this dissertation, we build heavily upon the Predictability, Computability, and Stability (PCS) framework and documentation for veridical (trustworthy) data science (Yu and Kumbier, 2020) to improve the reliability of our scientific conclusions. This framework advocates for the use of predictability as a reality check, computability as an important consideration in algorithmic design and data collection, and stability as a minimum requirement for reproducibility and interpretability in knowledge-seeking and decision-making. Moreover, it calls on the need for transparent documentation of decisions made throughout the data science pipeline.  \nIn Part I, we highlight two statistical machine learning methods, developed within the context of grounded biomedical problems and guided by the PCS framework. First, in Chapter 2, we investigate genetic and epistatic drivers of cardiac hypertrophy in hope of obtaining a more complete understanding of the disease architecture. To this end, we develop a data-driven recommendation system, named the low-signal signed iterative random forest (lo-siRF), to identify candidate genes and gene-gene interactions that are both predictive and stable across various model and data perturbations. We then phenotypically validate these genes and gene-gene interactions via gene-silencing experiments and investigate potential mechanistic explanations for the demonstrated epistases. This leads to a hypothesis in which the identified genes interact through mediating the variable binding of transcription factors that are essential for cardiac contractile function and metabolism. Second, the practical utility of random forests and interpretability tools, not only in the search for epistasis  \n2  \nbut in a wide range of scientific problems, motivates the need for reliable tree-based feature importance measures. In Chapter 3, we demonstrate that the mean decrease in impurity (MDI), arguably the most popular random forest feature importance measure, suffers from well-known biases including against highly-correlated and low-entropy features. To overcome these drawbacks, we develop a novel feature importance framework, MDI+, which leveragesa connection between MDI and the R2 value","cbCaim3QXy2tlEsH","https://ap.wps.com/l/cbCaim3QXy2tlEsH","pdf",16412783,1,228,"English","en",105,"# Overview\n## Part I: In-context development of interpretable machine learning methods guided by PCS\n## Part II: Responsible data science in real-world biomedical problems applying PCS\n## Part III: Open-source software and data\n# In-context development of interpretable machine learning methods guided by PCS\n## Low-signal iterative random forests (lo-siRF) for epistasis discovery","[{\"question\":\"What framework is used to improve the reliability of conclusions in this dissertation?\",\"answer\":\"The work builds on the Predictability, Computability, and Stability (PCS) framework for veridical (trustworthy) data science, using predictability as a reality check, computability as an algorithmic design consideration, and stability as a minimum requirement for reproducibility and interpretability.\"},{\"question\":\"How is lo-siRF used in Part I?\",\"answer\":\"lo-siRF is a data-driven recommendation system developed to identify candidate genes and gene-gene interactions related to cardiac hypertrophy, with validation via gene-silencing experiments and investigation of mechanistic explanations for observed epistasis.\"},{\"question\":\"What problem with random forest feature importance does the dissertation address?\",\"answer\":\"It shows that the mean decrease in impurity (MDI) suffers from biases against highly correlated and low-entropy features, then proposes MDI+ to improve reliability and stability of feature-importance rankings using a connection to linear regression R².\"}]","Statistical Machine Learning for Reliable Hypothesis Generation in Biomedical Problems - Thesis | PDF",1785934491,575,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"statistical-machine-learning-for-reliable-hypothesis-generation-in-biomedical-problems-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/statistical-machine-learning-for-reliable-hypothesis-generation-in-biomedical-problems-thesis/126732/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What framework is used to improve the reliability of conclusions in this dissertation?","Question",{"text":75,"@type":76},"The work builds on the Predictability, Computability, and Stability (PCS) framework for veridical (trustworthy) data science, using predictability as a reality check, computability as an algorithmic design consideration, and stability as a minimum requirement for reproducibility and interpretability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is lo-siRF used in Part I?",{"text":80,"@type":76},"lo-siRF is a data-driven recommendation system developed to identify candidate genes and gene-gene interactions related to cardiac hypertrophy, with validation via gene-silencing experiments and investigation of mechanistic explanations for observed epistasis.",{"name":82,"@type":73,"acceptedAnswer":83},"What problem with random forest feature importance does the dissertation address?",{"text":84,"@type":76},"It shows that the mean decrease in impurity (MDI) suffers from biases against highly correlated and low-entropy features, then proposes MDI+ to improve reliability and stability of feature-importance rankings using a connection to linear regression R².","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]