[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124306-en":3,"doc-seo-124306-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124306,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Safety, Robustness, and Interpretability in Machine Learning - Dissertation Abstract","Machine learning’s growing autonomy requires trustworthy behavior from increasingly large models. This dissertation targets three research priorities—safety, robustness, and interpretability—across reinforcement learning, imitation learning, and classifier security. It proposes a chance-constrained, model-predictive-control safety guide to refine base RL actions under user constraints, and uses structural causal models to mask causal confusion in imitation learning via initial-state interventions. It also develops certified robustness methods against adversarial perturbations, and introduces interpretability analyses for LLM-driven conversational search plus interpretable structural transport networks.","UC Berkeley  \nUC Berkeley Electronic Theses and Dissertations  \nTitle  \nSafety, Robustness, and Interpretability in Machine Learning  \nPermalink  \n[https://escholarship.org/uc/item/3915757j](https://escholarship.org/uc/item/3915757j)  \nISBN  \n9798288862991  \nAuthor  \nPfrommer, Samuel Ian  \nPublication Date  \n2025-05-24  \nPeer reviewed  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nSafety, Robustness, and Interpretability in Machine Learning  \nby  \nSamuel Ian Pfrommer  \nA dissertation submitted in partial satisfaction of the requirements for the degree of  \nDoctor of Philosophy  \nin  \nEngineering – Electrical Engineering and Computer Sciences  \nin the  \nGraduate Division  \nof the  \nUniversity of California, Berkeley  \nCommittee in charge:  \nProfessor Somayeh Sojoudi, Chair Professor Javad Lavaei Professor Venkatachalam Anantharam  \nSpring 2025  \nSafety, Robustness, and Interpretability in Machine Learning  \nCopyright 2025  \nby  \nSamuel Ian Pfrommer  \nAbstract  \nSafety, Robustness, and Interpretability in Machine Learning  \nby  \nSamuel Ian Pfrommer  \nDoctor of Philosophy in Engineering – Electrical Engineering and Computer Sciences  \nUniversity of California, Berkeley  \nProfessor Somayeh Sojoudi, Chair  \nMachine learning is poised to have a dramatic impact across many scientific, industrial, and social domains. While current Artificial Intelligence (AI) systems generally involve human supervision, future applications will demand significantly more autonomy. Such a transition will require us to trust the behavior of increasingly large models. This dissertation addresses three critical research areas towards this goal: safety, robustness, and interpretability.  \nWe first address safety concerns in Reinforcement Learning (RL) and Imitation Learning (IL) . While learned policies have achieved impressive performance, they often exhibit unsafe behavior due to training-time exploration and test-time environmental shifts. We introduce a model predictive control-based safety guide which refines the actions of a base RL policy, conditioned on user-provided constraints. With an appropriate optimization formulation and loss function, we show theoretically that the final base policy is provably safe at optimality. IL suffers from a distinct causal confusion safety concern, where spurious correlations between observations and expert actions can lead to unsafe behavior upon deployment. We leverage tools from Structural Causal Models (SCMs) to identify and mask problematic observations. Whereas previous work requires access to aqueryable expert or an expert reward function, our approach uses the typical ability of an experimenter to intervene on the initial state of an episode.  \nThe second part of this dissertation concerns robustifying machine learning classifiers against adversarial inputs. Classifiers are a critical component of many AI systems and have been shown to be highly sensitive to small input perturbations. We first extend randomized smoothing beyond traditional isotropic certification by projecting inputs into a data-manifold subspace, resulting in orders-of-magnitude improvements in certified volume. We then revisit the fundamental robustness problem by proposing asymmetric certification. This binary classification setting requires only certified robustness for one class, reflecting the fact that many real-world adversaries are strictly interested in producing false negatives. This more focused problem admits an interesting class of feature-convex architectures, which we leverage to provide efficient, deterministic, and closed-form certified radii.  \nThe third part of this dissertation discusses two distinct aspects of interpretability: how Large Language Models (LLMs) decide what to recommend to human users, and how we can build learned models which obey human-interpretable structures. We first analyze conversational search engines, in which we use LLMs to rank consu","cbCaiqdb2ijP7aDm","https://ap.wps.com/l/cbCaiqdb2ijP7aDm","pdf",2606848,1,157,"English","en",105,"# Introduction\n# I Safety\n## Safe Reinforcement Learning via Chance-Constrained Model Predictive Control","[{\"question\":\"What robustness methods are developed for classifiers against adversarial inputs?\",\"answer\":\"The dissertation extends randomized smoothing by projecting inputs into a data-manifold subspace for larger certified regions. It also proposes asymmetric certification for the binary setting, enabling efficient deterministic closed-form certified radii for one class.\"}]","Safety, Robustness, and Interpretability in Machine Learning - Dissertation Abstract | PDF",1785821499,396,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"safety-robustness-and-interpretability-in-machine-learning-dissertation-abstract","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/safety-robustness-and-interpretability-in-machine-learning-dissertation-abstract/124306/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What robustness methods are developed for classifiers against adversarial inputs?","Question",{"text":75,"@type":76},"The dissertation extends randomized smoothing by projecting inputs into a data-manifold subspace for larger certified regions. It also proposes asymmetric certification for the binary setting, enabling efficient deterministic closed-form certified radii for one class.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]