[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118475-en":3,"doc-seo-118475-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118475,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Machine Learning Methods in Biomedical Data for Precision Medicine - Thesis/dissertation","Recent advances in machine learning enable interrogation of diverse biomedical modalities, from genomic sequences to clinical records, yet progress is constrained by data heterogeneity, limited interpretability, and challenges in generalizing to robust real-world tools. This dissertation develops methods and benchmarks that strengthen predictive modeling for healthcare applications, emphasizing robustness, interpretability, and translational impact. It includes benchmarking variant calling, interpretable metagenomic taxonomic visualization, clinical risk modeling for urosepsis, and prediction of perinatal depression.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nMachine Learning Methods in Biomedical Data for Precision Medicine  \nPermalink  \n[https://escholarship.org/uc/item/1z66v3n4](https://escholarship.org/uc/item/1z66v3n4)  \nAuthor  \nSarwal, Varuni  \nPublication Date  \n2025  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nMachine Learning Methods in Biomedical Data for Precision Medicine  \nA dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Computer Science  \nby  \nVaruni Sarwal  \n© Copyright by  \nVaruni Sarwal  \n2025  \nABSTRACT OF THE DISSERTATION  \nMachine Learning Methods in Biomedical Data  \nfor Precision Medicine  \nby  \nVaruni Sarwal  \nDoctor of Philosophy in Computer Science  \nUniversity of California, Los Angeles, 2025  \nProfessor Eleazar Eskin, Chair  \nRecent advances in machine learning have created new opportunities to interrogate diverse biomedical data modalities, from genomic sequences to clinical records. However, challenges such as data heterogeneity, lack of interpretability, and limited generalizability have slowed the translation of these models into robust scientific tools and clinical applications. In this dissertation, I develop methods and benchmarks that enhance predictive modeling in real-world healthcare settings, with a focus on robustness, interpretability, and translational impact. I begin by benchmarking short-read structural variant callers on whole-genome sequencing data, identifying key limitations in existing methods which enables me to develop new methods that provide a balance of sensitivity and precision. I then present a method for taxonomic  \nvisualization that can enable interpretable analysis of metagenomic abundance profiles, helping researchers pinpoint identify differences across profilers where summary statistics fail.  \nIn the second half of the thesis, I shift focus to machine learning applications in clinical prediction. I develop a risk model for urosepsis using structured EHR data, achieving strong performance while highlighting risks with common treatment methods. Finally, I apply these insights to maternal health by developing a machine learning model to predict perinatal depression using structured health records. These models incorporate fairness assessments and identify both established and novel predictors across diverse patient populations. Finally, I evaluate the capabilities of large language models in bioinformatics through a comprehensive benchmarking effort that spans multiple scientific subdomains. Collectively, these contributions advance the design and evaluation of machine learning systems for genomics, microbiome research, and clinical informatics, with an emphasis on interpretability, fairness, and real-world utility.  \nThe dissertation of Varuni Sarwal is approved.  \nSerghei Mangul  \nLoes Marlein Olde Loohuis  \nJeffrey Nien-Jay Chiang  \nSriram Sankararaman  \nEleazar Eskin, Committee Chair  \nUniversity of California, Los Angeles  \n2025  \nTo my family—for giving me the faith that I could do anything To God—for giving me the strength to see it through  \nTABLE OF CONTENTS  \n1. Introduction  \n1.1 Biomedical Informatics  \n1.2 Structural Variation in the Human Genome  \n1.3 Microbiome Profiling and Metagenomic Analysis  \n1.4 Clinical Prediction from Electronic Health Records  \n1.5 Language Models and Benchmarking in Bioinformatics  \n2. VISTA: An integrated framework for structural variant discovery  \n2.1 Introduction  \n2.2 Methods  \n2.2.1 Preparing the datasets  \n2.2.2 VISTA algorithm  \n2.2.3 Comparing deletion inferred from WGS data with the gold standard  \n2.2.4 Downsample the WGS samples  \n2.2.5 VISTA Train-test experiments  \n2.2.6 Data availability  \n2.2.7 Code availability  \n2.2.8 Tool Availability  \n2.3 Results  \n2.3.1 VISTA: An integrated framework for structural variant","cbCaiaqHEy6SJb1m","https://ap.wps.com/l/cbCaiaqHEy6SJb1m","pdf",4610831,1,157,"English","en",105,"# Introduction\n## Biomedical Informatics\n## Structural Variation in the Human Genome\n## Microbiome Profiling and Metagenomic Analysis\n## Clinical Prediction from Electronic Health Records\n## Language Models and Benchmarking in Bioinformatics\n# VISTA: An integrated framework for structural variant discovery\n## Methods\n## Results\n## Discussion\n# TAMPA: interpretable analysis and visualization of metagenomics-based taxon abundance profiles\n## Introduction\n## Results\n## Discussion\n# Machine learning for the prediction of urosepsis using electronic health record data\n## Introduction\n## Methods\n## Results\n## Discussion\n# Early prediction and fairness evaluation of perinatal depression using EHR: A study of 18,000+ Pregnancies\n## Introduction\n## Methods\n## Results","[{\"question\":\"What core problems does the dissertation address in applying machine learning to biomedical data?\",\"answer\":\"It addresses data heterogeneity, limited interpretability, and limited generalizability that slow translation of machine learning models into reliable scientific and clinical tools.\"},{\"question\":\"How does the work evaluate methods for genomics applications?\",\"answer\":\"It benchmarks short-read structural variant callers on whole-genome sequencing data, identifies limitations, and develops approaches that balance sensitivity and precision.\"},{\"question\":\"Which clinical prediction tasks are modeled in the second half of the thesis?\",\"answer\":\"The thesis develops a risk model for urosepsis using structured EHR data and a model to predict perinatal depression from structured health records, including fairness assessments and identification of predictors across patient populations.\"}]","Machine Learning Methods in Biomedical Data for Precision Medicine - Thesis/dissertation | PDF",1785683789,396,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"machine-learning-methods-in-biomedical-data-for-precision-medicine-thesisdissertation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-methods-in-biomedical-data-for-precision-medicine-thesisdissertation/118475/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What core problems does the dissertation address in applying machine learning to biomedical data?","Question",{"text":76,"@type":77},"It addresses data heterogeneity, limited interpretability, and limited generalizability that slow translation of machine learning models into reliable scientific and clinical tools.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the work evaluate methods for genomics applications?",{"text":81,"@type":77},"It benchmarks short-read structural variant callers on whole-genome sequencing data, identifies limitations, and develops approaches that balance sensitivity and precision.",{"name":83,"@type":74,"acceptedAnswer":84},"Which clinical prediction tasks are modeled in the second half of the thesis?",{"text":85,"@type":77},"The thesis develops a risk model for urosepsis using structured EHR data and a model to predict perinatal depression from structured health records, including fairness assessments and identification of predictors across patient populations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]