[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124842-en":3,"doc-seo-124842-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124842,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Robust Modeling through Causal Priors and Data Purification in Machine Learning","Machine learning’s widespread use, especially deep learning, demands robust training that improves generalization and security under incomplete data, distribution shifts, and adversarial attacks. This thesis develops two contribution tracks for robust modeling using causal priors and data purification with generative models (VAE, EBM, DDPM) on image datasets, including counterfactual latent-space generation, hypothesis comparison via causal structural assumptions, and universal purification defenses that mitigate data poisoning while preserving natural accuracy and enabling further applications.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nRobust Modeling through Causal Priors and Data Purification in Machine Learning  \nPermalink  \n[https://escholarship.org/uc/item/98f7s81c](https://escholarship.org/uc/item/98f7s81c)  \nAuthor  \nBhat, Sunay Gajanan  \nPublication Date  \n2024  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nRobust Modeling through Causal Priorsand Data Purification in Machine Learning  \nA dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Electrical and Computer Engineering  \nby  \nSunay Gajanan Bhat  \n2024  \n© Copyright by Sunay Gajanan Bhat 2024  \nABSTRACT OF THE DISSERTATION  \nRobust Modeling through Causal Priors  \nand Data Purification in Machine Learning  \nby  \nSunay Gajanan Bhat  \nDoctor of Philosophy in Electrical and Computer Engineering University of California, Los Angeles, 2024  \nProfessor Gregory J. Pottie, Chair  \nThe continued success and ubiquity of machine learning techniques, particularly Deep Learning, have necessitated research in robust model training to enhance generalization capabilities and security against incomplete data, distributional shifts, and adversarial attacks. This thesis presents two primary sets of contributions to robust modeling in machine learning through the use of causal priors and data purification with generative models such as the Variational Autoencoder (VAE), Energy-Based Model (EBM), and Denoising Diffusion Probabilistic Model (DDPM), focusing on image datasets. In the first set of contributions, we use structural causal priors in the latent spaces of VAEs. Initially, we demonstrate counterfactual synthetic data generation outside the training data distribution. This technique allows for the creation of diverse and novel data points, which is critical to enhancing model robustness and generalization capabilities. We utilize a similar VAE architecture to compare causal structural (graphical) hypotheses, showing that the fit of generated data from various hypotheses on distributionally shifted test data is an effective method for hypothesis comparison. Additionally, we explore using augmentations in the latent space of a VAE  \nas an efficient and effective way to generate realistic augmented data. The second set of contributions focuses on data purification using EBMs and DDPMs. We propose a framework of universal data purification methods to defend against train-time data poisoning attacks. This framework utilizes stochastic transforms realized via iterative Langevin dynamics of EBMs, DDPMs, or both, to purify poisoned data with minimal impact on classifier generalization. Our specially trained EBMs and DDPMs provide state-of-the-art defense against various poisoning attacks while preserving natural accuracy. Preprocessing data with these techniques pushes poisoned images into the natural, clean image manifold, effectively neutralizing adversarial perturbations. The framework achieves state-of-the-art performance without needing attack or classifier-specific information, even when the generative models are trained on poisoned or distributionally shifted data. Beyond defense against data poisoning, our framework also shows promise in applications such as the degradation and removal of unwanted intellectual property. The flexibility and generality of these data purification techniques represent a significant step forward in the adversarial model training paradigm. All of these methods enable new perspectives and approaches to robust machine learning, advancing an essential field in artificial intelligence research.  \nThe dissertation of Sunay Gajanan Bhat is approved.  \nJonathan Chau-Yan Kao  \nAchuta Kadambi  \nLieven Vandenberghe Gregory J. Pottie, Committee Chair  \nUniversity of California, Los Angeles  \n2024  \nTo my loving wife and better half, Cindy, to my s","cbCaiuinifyQvKbS","https://ap.wps.com/l/cbCaiuinifyQvKbS","pdf",17168052,1,218,"English","en",105,"# Thesis Overview\n## Motivation\n## Research Objectives and Contributions\n## Thesis organization\n# Background\n## Causality and Causal Machine Learning\n## Deep Learning\n## Preliminary Work on Structural Causal Neural Networks\n## Generative Modeling and Causality\n## Adversarial Training and Data Purification\n# De-Biasing Generative Models using Counterfactual Methods\n## Introduction\n## Related Work\n## Background\n## Counterfactuals and Interventions\n## Counterfactual Models","[{\"question\":\"What does the thesis focus on in robust machine learning training?\",\"answer\":\"It focuses on improving generalization and security against incomplete data, distributional shifts, and adversarial attacks through causal priors and data purification methods.\"},{\"question\":\"How are causal priors used in the first set of contributions?\",\"answer\":\"Causal structural priors are applied in the latent spaces of VAEs to enable counterfactual synthetic data generation and to compare causal graphical hypotheses using performance on distribution-shifted test data.\"},{\"question\":\"How does the second set of contributions defend against data poisoning?\",\"answer\":\"It proposes universal data purification using EBMs and DDPMs with stochastic transforms realized via iterative Langevin dynamics, pushing poisoned samples toward a clean image manifold while preserving classification generalization.\"}]","Robust Modeling through Causal Priors and Data Purification in Machine Learning | PDF",1785894947,549,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robust-modeling-through-causal-priors-and-data-purification-in-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/robust-modeling-through-causal-priors-and-data-purification-in-machine-learning/124842/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the thesis focus on in robust machine learning training?","Question",{"text":75,"@type":76},"It focuses on improving generalization and security against incomplete data, distributional shifts, and adversarial attacks through causal priors and data purification methods.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are causal priors used in the first set of contributions?",{"text":80,"@type":76},"Causal structural priors are applied in the latent spaces of VAEs to enable counterfactual synthetic data generation and to compare causal graphical hypotheses using performance on distribution-shifted test data.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the second set of contributions defend against data poisoning?",{"text":84,"@type":76},"It proposes universal data purification using EBMs and DDPMs with stochastic transforms realized via iterative Langevin dynamics, pushing poisoned samples toward a clean image manifold while preserving classification generalization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]