[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119149-en":3,"doc-seo-119149-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119149,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Distributional Robustness and Machine Learning Methods for Stochastic Decision-Making - Doctoral Thesis","Data-driven decision-making increasingly influences classification, prediction, optimization, and resource allocation, yet standard machine learning models struggle with uncertainty from future data or new application domains. This thesis studies how distributionally robust optimization can be integrated with modern learning to improve reliability and reduce performance risk under distribution shifts. It provides new theoretical results and strong experiments across three themes: generalization bounds for DRO, prediction-error queue scheduling, and tabular classification under |·| shifts, leveraging LLM embeddings to boost few-shot target performance.","Distributional robustness and machine learning methods for stochastic decision-making  \nYibo Zeng  \nSubmitted in partial fulfillment of the  \nrequirements for the degree of  \nDoctor of Philosophy  \nunder the Executive Committee  \nof the Graduate School of Arts and Sciences  \nCOLUMBIA UNIVERSITY  \n© 2024 Yibo Zeng All Rights Reserved  \nAbstract  \nDistributional Robustness and Machine Learning Methods for Stochastic Decision-Making  \nYibo Zeng  \nRecently, data-driven methods are increasingly shaping decision-making across domains such as classification, prediction, optimization, and resource allocation. While machine learning advancements have been pivotal, current models can be insufficient to handle unseen uncertainties from future data or new application domains, leading to unreliable decisions and performance risks. This thesis explores how distributionally robust methods can be combined with modern machine learning techniques to ensure more reliable decision-making. We develop new theoretical results and achieve state-of-the-art empirical results in three areas: generalization bounds in machine learning, queue scheduling with prediction errors, and tabular classification under 􀀮 | 􀀭 shifts. The work is organized into three chapters. In Chapter 2, we derive novel generalization bounds for distributionally robust optimization (DRO) . Our analysis implies generalization bounds whose dependence on the hypothesis class appears the minimal possible: The bound depends solely on the true loss function, independent of any other candidates in the hypothesis class. To our best knowledge, it is the first generalization bound of this type in the literature. Chapter 3 investigates optimal scheduling in service systems with prediction errors. We develop a near-optimal  \nindex-based policy that incorporates predicted class information. Our results guide model selection  \nwith a focus on downstream queueing performance and offer insights into designing queueing systems with AI-based triage. In Chapter 4, we study tabular data classification under 􀀮 | 􀀭 shift. We not only build a large-scale testbed, but also demonstrate that large language model (LLM) embeddings significantly improve classification performance, even with few labeled samples from the target domain.  \nTable of Contents  \nAcknowledgments ........................................ vii  \nDedication ............................................ viii  \nChapter 1: Introduction .................................... 1  \nChapter 2: Generalization Power of Distributionally Robust Optimization: From Worst-Case Analysis to Minimal Hypothesis Class Dependence ............... 3  \n2.1 Introduction ...................................... 3  \n2.2 Related Work and Comparisons ............................ 7  \n2.2.1 Absolute Bounds on Expected Loss ..................... 7  \n2.2.2 Variability Regularization ........................... 9  \n2.2.3 Risk Aversion ................................. 11  \n2.3 General Results .................................... 11  \n2.3.1 A Beginning Bound .............................. 11  \n2.3.2 Specialization to G-IPM DRO ........................ 13  \n2.3.3 Strengths and Limitations Compared with ERM ............... 15  \n2.4 Specialization to MMD DRO ............................. 17  \n2.4.1 Comparison with [21] ............................. 18  \n2.4.2 Comparison with RKHS Regularized ERM ................. 19  \n2.4.3 Extension to ℓ( 􀀵 ∗ , ·) ∉ H ........................... 19  \n2.5 Specialization to 1-Wasserstein DRO ......................... 21  \n2.6 Specialization to 􀁱-divergence DRO ......................... 22  \n2.6.1 General Bounds for 􀁱-divergence DRO ................... 23  \n2.6.2 Specialization to Cressie-Read divergence .................. 24  \n2.6.3 Specialization to 􀁪2-divergence ....................... 26  \n2.7 Extensions to Distributional Shift Settings ...................... 27  \n2.8 Numerical Experiments ................................ 31  \n2.8.1 A Special Cas","cbCaiiSBPkqjX7Dx","https://ap.wps.com/l/cbCaiiSBPkqjX7Dx","pdf",2831392,1,223,"English","en",105,"# Chapter 1: Introduction\n# Chapter 2: Generalization Power of Distributionally Robust Optimization: From Worst-Case Analysis to Minimal Hypothesis Class Dependence\n## 2.2 Related Work and Comparisons\n## 2.3 General Results\n## 2.4 Specialization to MMD DRO\n## 2.5 Specialization to 1-Wasserstein DRO\n## 2.6 Specialization to f-divergence DRO\n## 2.7 Extensions to Distributional Shift Settings\n## 2.8 Numerical Experiments\n# Chapter 3: Design and Scheduling of an AI-based Queueing System\n## 3.1 Introduction\n## 3.2 Model\n## 3.3 Lower bound on queueing cost\n## 3.4 Heavy-traffic optimality of the P^2 φ-rule\n## 3.6 Model selection based on queueing cost\n## 3.7 Design of an AI-based triage system\n## 3.8 Discussion\n# Chapter 4: LLM Embeddings Improve Test-time Adaptation to Tabular |·|-Shifts\n## 4.1 Introduction\n## 4.2 Methods\n## 4.3 Numerical Experiments\n## 4.4 Discussion","[{\"question\":\"Why do current machine learning models struggle with stochastic decision-making under future uncertainty?\",\"answer\":\"They may not adequately handle unseen uncertainties caused by future data or new application domains, which can lead to unreliable decisions and performance risk.\"},{\"question\":\"What is the key theoretical contribution of Chapter 2?\",\"answer\":\"It derives distributionally robust optimization generalization bounds where dependence on the hypothesis class is minimized, relying solely on the true loss function.\"},{\"question\":\"How does Chapter 3 use prediction information for queue scheduling?\",\"answer\":\"It develops a near-optimal index-based policy that incorporates predicted class information and guides model selection based on downstream queueing performance.\"},{\"question\":\"How do LLM embeddings help in Chapter 4 for tabular classification under distribution shifts?\",\"answer\":\"They significantly improve classification performance on the target domain, even when only few labeled samples are available, enabling better test-time adaptation.\"}]","Distributional Robustness and Machine Learning Methods for Stochastic Decision-Making - Doctoral Thesis | PDF",1785722743,562,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"distributional-robustness-and-machine-learning-methods-for-stochastic-decision-making-doctoral-thesis","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/distributional-robustness-and-machine-learning-methods-for-stochastic-decision-making-doctoral-thesis/119149/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"Why do current machine learning models struggle with stochastic decision-making under future uncertainty?","Question",{"text":75,"@type":76},"They may not adequately handle unseen uncertainties caused by future data or new application domains, which can lead to unreliable decisions and performance risk.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key theoretical contribution of Chapter 2?",{"text":80,"@type":76},"It derives distributionally robust optimization generalization bounds where dependence on the hypothesis class is minimized, relying solely on the true loss function.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Chapter 3 use prediction information for queue scheduling?",{"text":84,"@type":76},"It develops a near-optimal index-based policy that incorporates predicted class information and guides model selection based on downstream queueing performance.",{"name":86,"@type":73,"acceptedAnswer":87},"How do LLM embeddings help in Chapter 4 for tabular classification under distribution shifts?",{"text":88,"@type":76},"They significantly improve classification performance on the target domain, even when only few labeled samples are available, enabling better test-time adaptation.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]