[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117001-en":3,"doc-seo-117001-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117001,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Toward Trustworthy Scientific Inquiry and Design with Machine Learning","Machine-learning systems are increasingly used in science for rapid predictions and for discovering knowledge and designing new biomolecules, yet their errors raise a central trust problem. This dissertation develops methods to ensure trust in both designed biomolecules and resulting scientific conclusions. It addresses machine-learning-based design under distribution shift and constructs statistically valid confidence sets for properties of designed objects. It further introduces prediction-powered inference, treating model outputs as data to form valid confidence sets for scientific quantities.","UC Berkeley  \nUC Berkeley Electronic Theses and Dissertations  \nTitle  \nToward Trustworthy Scientific Inquiry and Design with Machine Learning  \nPermalink  \n[https://escholarship.org/uc/item/67h0m1qj](https://escholarship.org/uc/item/67h0m1qj)  \nAuthor  \nWong-Fannjiang, Clara  \nPublication Date  \n2023  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nToward Trustworthy Scientific Inquiry and Design with Machine Learning  \nBy  \nClara Wong-Fannjiang  \nA dissertation submitted in partial satisfaction of the requirements for the degree of  \nDoctor of Philosophy  \nin  \nEngineering – Electrical Engineering and Computer Sciences  \nin the  \nGraduate Division  \nof the  \nUniversity of California, Berkeley  \nCommittee in charge:  \nProfessor Michael I. Jordan, Co-chair Professor Jennifer Listgarten, Co-chair  \nProfessor David V. Schaffer  \nSummer 2023  \nToward Trustworthy Scientific Inquiry and Design with Machine Learning  \nCopyright 2023  \nby  \nClara Wong-Fannjiang  \n1  \nAbstract  \nToward Trustworthy Scientific Inquiry and Design with Machine Learning  \nby  \nClara Wong-Fannjiang  \nDoctor of Philosophy in Engineering – Electrical Engineering and Computer Sciences  \nUniversity of California, Berkeley  \nProfessor Michael I. Jordan, Co-chair  \nProfessor Jennifer Listgarten, Co-chair  \nThe last decade has witnessed rapid development and deployment of machine-learning systems across science. Such systems can supply predictions about scientific phenomena far more quickly and cheaply than gold-standard experiments, and are being used in efforts to both discover scientific knowledge and design new biomolecules. However, an important question remains unanswered: since machine-learning systems make errors, how can we use them in a trustworthy way for scientific discovery and design? This dissertation takes steps toward helping to ensure that the biomolecules we design and the scientific conclusions we draw using machine learning can be trusted.  \nWe begin in the setting of machine learning-based design. The goal in this setting is to propose novel objects such as proteins, small molecules, or materials with desired properties, in a way that is guided by machine-learning models of such properties. Toward addressing model trustworthiness for design, we propose (i) a method for learning models that accounts for the distribution shifts inherent to design, and (ii) a method for constructing statistically valid confidence sets for the properties of objects designed using machine learning.  \nFinally, we examine the trustworthy use of machine learning for drawing scientific conclusions. In particular, we consider the increasingly relevant setting of treating predictions made by machine-learning systems as “data” in estimating quantities of scientific interest. We propose prediction-powered inference, a novel statistical framework for constructing valid confidence sets in this setting, which enables researchers to incorporate evidence from machine-learning systems into their scientific inquiry in a standardized and principled way.  \ni  \nIn loving memory of Pei-rong Wang, who hoped his children’s children could chase dreams.  \nii  \nTest all things; hold fast what is good.  \n– 1 Thessalonians 5:21  \niii  \nContents  \nContents iii  \nList of Figures v  \nList of Tables vii  \n1 Introduction 1  \n2 Autofocused surrogates for design 5  \n2.1 Surrogates for design ............................... 5  \n2.2 Model-based optimization for design ...................... 6  \n2.3 Autofocused surrogates for design ........................ 7  \n2.4 Related work ................................... 11  \n2.5 Experiments .................................... 12  \n2.6 Discussion ..................................... 17  \n2.7 Pseudocode .................................... 18  \n2.8 Proofs and derivations .............................. 18  \n2.9 Experimental details ..............................","cbCairVRPc0WZriI","https://ap.wps.com/l/cbCairVRPc0WZriI","pdf",13265200,1,116,"English","en",105,"# Introduction\n# Autofocused surrogates for design\n## Surrogates for design\n## Model-based optimization for design\n## Autofocused surrogates for design\n## Related work\n## Experiments\n## Discussion\n## Pseudocode\n## Proofs and derivations\n## Experimental details\n# Conformal prediction under feedback covariate shift for biomolecular design\n## Uncertainty quantification under feedback loops\n## Conformal prediction under feedback covariate shift\n## Simulated protein design experiments\n## Discussion\n## Proofs\n## Data splitting\n## Efficient algorithms for full conformal prediction\n## Experimental details and additional results\n# Prediction-powered inference\n## Predictions as data for scientific inquiry\n## Main theory: estimands that minimize convex objectives\n## Algorithms\n## Applications in proteomics and genomics\n## Proofs\n## Experimental details","[{\"question\":\"Why is trustworthy use of machine learning important for scientific discovery and design?\",\"answer\":\"Machine-learning systems can produce predictions and support biomolecule design, but they make errors. The dissertation addresses how to use those outputs in a trustworthy way for both designed biomolecules and downstream scientific conclusions.\"},{\"question\":\"What trustworthiness methods are proposed for machine-learning-based design?\",\"answer\":\"The work proposes (i) a method for learning models that accounts for distribution shifts inherent to design, and (ii) a method to construct statistically valid confidence sets for the properties of designed objects.\"},{\"question\":\"How does prediction-powered inference improve scientific conclusions?\",\"answer\":\"It provides a statistical framework that treats predictions from machine-learning systems as “data” when estimating scientific quantities. This enables valid confidence sets in a standardized and principled way.\"}]","Toward Trustworthy Scientific Inquiry and Design with Machine Learning | PDF",1785673032,292,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"toward-trustworthy-scientific-inquiry-and-design-with-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/toward-trustworthy-scientific-inquiry-and-design-with-machine-learning/117001/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is trustworthy use of machine learning important for scientific discovery and design?","Question",{"text":75,"@type":76},"Machine-learning systems can produce predictions and support biomolecule design, but they make errors. The dissertation addresses how to use those outputs in a trustworthy way for both designed biomolecules and downstream scientific conclusions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What trustworthiness methods are proposed for machine-learning-based design?",{"text":80,"@type":76},"The work proposes (i) a method for learning models that accounts for distribution shifts inherent to design, and (ii) a method to construct statistically valid confidence sets for the properties of designed objects.",{"name":82,"@type":73,"acceptedAnswer":83},"How does prediction-powered inference improve scientific conclusions?",{"text":84,"@type":76},"It provides a statistical framework that treats predictions from machine-learning systems as “data” when estimating scientific quantities. This enables valid confidence sets in a standardized and principled way.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]