[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116911-en":3,"doc-seo-116911-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116911,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Theory-Guided Algorithm Design for Scalable Machine Learning - Dissertation","This dissertation develops scalable machine learning algorithms grounded in mathematical theory. The work addresses two scalability-critical directions: fair machine learning under individual fairness in both single-model and decoupled-model settings with constrained data labeling budgets, and randomized feature representations via a model-agnostic framework offering computational efficiency and provable performance guarantees. It further advances scalable estimation of Kernel matrix spectral norm using sketching techniques, including theoretical estimation error bounds and empirical efficiency under time constraints.","UNIVERSITY OF OKLAHOMA  \nGRADUATE COLLEGE  \nTHEORY-GUIDED ALGORITHM DESIGN  \nFOR SCALABLE MACHINE LEARNING  \nA DISSERTATION  \nSUBMITTED TO THE GRADUATE FACULTY in partial fulfillment of the requirements for the Degree of Doctor of Philosophy  \nBy  \nYITING CAO Norman, Oklahoma May 2023  \nTHEORY-GUIDED ALGORITHM DESIGN  \nFOR SCALABLE MACHINE LEARNING  \nA DISSERTATION APPROVED FOR  \nTHE SCHOOL OF COMPUTER SCIENCE  \nBY THE COMMITTEE CONSISTING OF  \nDr. Chao Lan  \nDr. Dean Hougen  \nDr. Dimitris Diochnos  \nDr. Christian Remling  \n©Copyright by YITING CAO 2023 All Rights Reserved.  \nAbstract  \nMy thesis focuses on designing scalable machine learning algorithms leveraging theoretical advances in mathematics. In particular, I investigate two directions where scalability plays an important role: fair machine learning and randomized feature representations. In fair machine learning, my research concentrates on achieving individual fairness in the single model and decoupled model settings with minimum data labeling budgets. For randomized feature representations, I propose a model-agnostic framework for designing computationally efficient randomized machine learning algorithms with provable performance guarantees, which demonstrates that it is not necessary for individual models to be weakly trained before they are optimally ensembled. Furthermore, I also contribute to the scalable estimation of Kernel matrix spectral norm. Specifically, I propose to apply sketching techniques to efficiently estimate the spectral norm, theoretically derive the estimation error and empirically demonstrate the estimation efficiency in a time-constrained setting.  \nAcknowledgments  \nFirst and foremost, I want to thank my advisor Dr. Chao Lan for always believing in my potential and consistently supporting my career choices, for example, consulting another faculty for opinion on certain research topics, having summer internship at Quora, and graduating early. Next, I want to thank my committee members for helpful discussions during our meetings. I also want to thank the School of Computer Science at University of Oklahoma for providing support throughout my years here. Finally, I want to thank my parents for supporting my educational journey. I will not become who I am today without them.  \nAll my papers mentioned in this thesis are joint work with Dr. Chao Lan.  \nContents  \n1 Introduction 1  \n1.1 Fair Machine Learning ..................... 1  \n1.2 Randomized Feature Representations ............. 2  \n1.3 Outline .............................. 3  \n2 Preliminaries 4  \n2.1 Tools from Matrix Analysis and High Dimensional Probability 4  \n2.2 Tools from Previous Research ................. 5  \nI Active Fair Learning 7  \n3 Background on Active Fair Learning 8  \n3.1 Introduction ........................... 8  \n3.2 Related Work .......................... 9  \n3.2.1 Fair Learning ...................... 9  \n3.2.2 Active Learning ..................... 10  \n3.2.3 Fairness in Active Learning .............. 10  \n4 Active Approximately Metric-Fair (AMF) Learning 12  \n4.1 Active AMF Learning ...................... 12  \n4.1.1 Approximate Metric-Fairness (AMF) ......... 12  \n4.1.2 Sample Complexity of Active AMF Learning ..... 15  \n4.1.3 The Counter AMF Coefficient ............. 17  \n4.2 Experiments ........................... 20  \n4.2.1 Implementation Issues ................. 20  \n4.2.2 Empirical Results .................... 21  \n4.3 Conclusion ............................ 24  \n4.4 Supplementary Material .................... 24  \n5 Fairness-Aware Active Learning for Decoupled Model 31  \n5.1 Proposed Algorithm ...................... 31  \n5.2 Theoretical Analysis ...................... 31  \n5.2.1 Notations and Definitions ............... 32  \n5.2.2 Main Theoretical Results ................ 34  \n5.2.3 Impact of D-FA2 L on Model Accuracy ........ 37  \n5.3 Experiments ........................... 38  \n5.3.1 Data Preparation .................... 38  \n5.3.2 Experiment Design ...........","cbCaiiXz4zZho7Qp","https://ap.wps.com/l/cbCaiiXz4zZho7Qp","pdf",1980127,1,107,"English","en",105,"# Introduction\n## Fair Machine Learning\n## Randomized Feature Representations\n## Outline\n# Preliminaries\n## Tools from Matrix Analysis and High Dimensional Probability\n## Tools from Previous Research\n# Active Fair Learning\n## Background on Active Fair Learning\n## Related Work\n# Active Approximately Metric-Fair (AMF) Learning\n## Active AMF Learning\n## Experiments\n## Conclusion\n## Supplementary Material\n# Fairness-Aware Active Learning for Decoupled Model\n## Proposed Algorithm\n## Theoretical Analysis\n## Experiments\n## Conclusion\n## Proof\n# Randomized Machine Learning Methods\n## A Model-Agnostic Randomized Learning Framework based on Random Hypothesis Subspace Sampling\n## SpectralSketches: Scaling Up Spectral Norm Estimation through Sketchings","[{\"question\":\"What are the two main research directions for scalability in this dissertation?\",\"answer\":\"The dissertation focuses on fair machine learning and randomized feature representations, both designed to remain scalable under practical constraints such as labeling budgets and computational efficiency requirements.\"},{\"question\":\"How does the work approach individual fairness under labeling constraints?\",\"answer\":\"It targets individual fairness in single-model and decoupled-model settings while minimizing data labeling budgets, including active variants with theoretical sample-complexity style analysis.\"},{\"question\":\"How is kernel matrix spectral norm estimation made scalable?\",\"answer\":\"The dissertation applies sketching techniques to estimate the spectral norm efficiently, provides theoretical estimation error bounds, and validates estimation efficiency in time-constrained experiments.\"}]","Theory-Guided Algorithm Design for Scalable Machine Learning - Dissertation | PDF",1785672458,270,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"theory-guided-algorithm-design-for-scalable-machine-learning-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/theory-guided-algorithm-design-for-scalable-machine-learning-dissertation/116911/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two main research directions for scalability in this dissertation?","Question",{"text":75,"@type":76},"The dissertation focuses on fair machine learning and randomized feature representations, both designed to remain scalable under practical constraints such as labeling budgets and computational efficiency requirements.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the work approach individual fairness under labeling constraints?",{"text":80,"@type":76},"It targets individual fairness in single-model and decoupled-model settings while minimizing data labeling budgets, including active variants with theoretical sample-complexity style analysis.",{"name":82,"@type":73,"acceptedAnswer":83},"How is kernel matrix spectral norm estimation made scalable?",{"text":84,"@type":76},"The dissertation applies sketching techniques to estimate the spectral norm efficiently, provides theoretical estimation error bounds, and validates estimation efficiency in time-constrained experiments.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]