[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120346-en":3,"doc-seo-120346-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120346,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Machine Learning Approaches to Empower Novel Discoveries from Biomedical Data - research thesis","Unprecedented clinical data availability drives the need for advanced machine learning methods that improve multiple aspects of healthcare, while addressing statistical, computational, and conceptual hurdles distinct from related domains. This dissertation presents computational approaches for large-scale biomedical data, including deep learning–based imputation for jointly analyzing genomic and sparse phenotypic data in biobanks. It further develops methods for imputing dense genomic information, enhancing interpretability via latent variable models, transferring knowledge from pretrained language models to disease code assignment, and generating compressed representations of medical imaging volumes to enable downstream discoveries.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nMachine Learning Approaches to Empower Novel Discoveries from Biomedical Data  \nPermalink  \n[https://escholarship.org/uc/item/5rw5n2d3](https://escholarship.org/uc/item/5rw5n2d3)  \nAuthor  \nAn, Ulzee  \nPublication Date  \n2025  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nMachine Learning Approaches to Empower Novel Discoveries from Biomedical Data  \nA dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Computer Science  \nby  \nUlzee An  \n© Copyright by  \nUlzee An  \n2025  \nABSTRACT OF THE DISSERTATION  \nMachine Learning Approaches to Empower Novel Discoveries  \nfrom Biomedical Data  \nby  \nUlzee An  \nDoctor of Philosophy in Computer Science  \nUniversity of California, Los Angeles, 2025  \nProfessor Wei Wang, Committee Co-Chair  \nProfessor Sriram Sankararaman, Committee Co-Chair  \nThe unprecedented availability of clinical data for machine learning research has motivated the exploration of state-of-the-art methods to improve various aspects of healthcare. However, developing machine learning algorithms to understand human health presents unique statistical, computational, and conceptual challenges not typically encountered in adjacent domains. In this dissertation, I propose several advanced computational methods to address emerging challenges in studying largescale biomedical data. To overcome challenges in jointly analyzing genomic and sparse phenotypic data, I first propose a deep learning–based imputation method to approximate missing assessmentsin biobanks, enabling the joint analysis of genomic and phenotypic data at a scale that was previously not explored. I then explore the effectiveness of an approach to impute high-density genomic data, which demonstrates favorable properties compared to existing methods based on Hidden-Markov Models. To improve the interpretability of concepts learned by such models, I propose both a statistical and a deep learning–based latent variable model that identifies the proportion of unknown subtypes present in data (applied to microbiome and longitudinal patient health data) . In exploring the latent space of these models, I further propose an approach to transfer prior knowledge embedded in pretrained language models to improve a disease code assignment model. Finally, I propose a novel method to obtain compressed representations of medical volumes (e.g., MRIs) that overcomes logistical challenges in studying such modalities and empowers new findings in downstream analysis.  \nThe dissertation of Ulzee An is approved.  \nEran Halperin Noah A. Zaitlen Quanquan Gu Wei Wang, Co-Chair Sriram Sankararaman, Co-Chair  \nUniversity of California, Los Angeles 2025  \nThis thesis is dedicated to my parents for their sacrifices, love, and support.  \nTABLE OF CONTENTS  \n1 Introduction 1  \n2 Overcoming sparse observations in EHR using deep learning-based imputation 5  \n2.1 Background ......................................... 5  \n2.2 Methods ........................................... 8  \n2.2.1 AutoComplete: Denoising autoencoder approach to target structured missingness ......................................... 8  \n2.2.2 Copy Masking ................................... 11  \n2.2.3 Experiment setup .................................. 12  \n2.3 Results ............................................ 14  \n2.3.1 AutoComplete significantly improved imputation accuracy ........... 14  \n2.3.2 Imputed phenotypes lead to replicable genomic discoveries ........... 16  \n2.4 Discussion .......................................... 22  \n2.5 Appendix .......................................... 23  \n2.5.1 Evaluation of runtime ............................... 23  \n2.5.2 Change in genomic analysis after accounting for uncertainty .......... 26  \n2.5.3 Tests in Smaller Scales ..............","cbCaim0OjfYVUl1M","https://ap.wps.com/l/cbCaim0OjfYVUl1M","pdf",14490458,1,186,"English","en",105,"# 1 Introduction\n## 2 Overcoming sparse observations in EHR using deep learning-based imputation\n## 3 Scaling deep learning architectures to whole-genome variant imputation\n## 4 Scaling microbial source tracking to large environments","[{\"question\":\"What motivates this dissertation on machine learning for healthcare?\",\"answer\":\"The rapid availability of clinical data motivates state-of-the-art methods to improve healthcare, while recognizing unique challenges in understanding human health from biomedical data.\"},{\"question\":\"How does the dissertation address missing or sparse biomedical observations?\",\"answer\":\"It proposes deep learning–based imputation methods, including a model for approximating missing assessments in biobanks and approaches for imputing high-density genomic data.\"},{\"question\":\"What techniques improve interpretability and downstream usefulness of learned models?\",\"answer\":\"It introduces statistical and deep latent variable models to identify proportions of unknown subtypes, uses knowledge transfer from pretrained language models for disease code assignment, and proposes compressed representations for medical volumes to support downstream analysis.\"}]","Machine Learning Approaches to Empower Novel Discoveries from Biomedical Data - research thesis | PDF",1785729592,469,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-approaches-to-empower-novel-discoveries-from-biomedical-data-research-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-approaches-to-empower-novel-discoveries-from-biomedical-data-research-thesis/120346/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What motivates this dissertation on machine learning for healthcare?","Question",{"text":75,"@type":76},"The rapid availability of clinical data motivates state-of-the-art methods to improve healthcare, while recognizing unique challenges in understanding human health from biomedical data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the dissertation address missing or sparse biomedical observations?",{"text":80,"@type":76},"It proposes deep learning–based imputation methods, including a model for approximating missing assessments in biobanks and approaches for imputing high-density genomic data.",{"name":82,"@type":73,"acceptedAnswer":83},"What techniques improve interpretability and downstream usefulness of learned models?",{"text":84,"@type":76},"It introduces statistical and deep latent variable models to identify proportions of unknown subtypes, uses knowledge transfer from pretrained language models for disease code assignment, and proposes compressed representations for medical volumes to support downstream analysis.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]