[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121729-en":3,"doc-seo-121729-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121729,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Integrating Machine Learning and Optimization for Problems in Contextual Decision-Making and Dynamic Learning - Thesis","This thesis studies the intersection of optimization and machine learning for decision-making in contextual settings. It proposes a policy-evaluation framework for personalized pricing using optimization to assess new strategies against worst-case revenue functions. It then compares estimate-then-optimize versus integrated estimation–optimization under first-order stochastic dominance, showing when traditional predict-then-optimize prevails or reverses under model misspecification. The work further introduces Implicit Two-Tower reinforcement learning policies, and analyzes active learning limits for nonparametric classification under margin notions.","Integrating Machine Learning and Optimization for Problems in Contextual Decision-Making and  \nDynamic Learning.  \nYunfan Zhao  \nSubmitted in partial fulfillment of the  \nrequirements for the degree of  \nDoctor of Philosophy  \nunder the Executive Committee  \nof the Graduate School of Arts and Sciences  \nCOLUMBIA UNIVERSITY  \n© 2023 Yunfan Zhao  \nAll Rights Reserved  \nAbstract  \nIntegrating Machine Learning and Optimization for Problems in Contextual Decision-Making and  \nDynamic Learning  \nYunfan Zhao  \nIn this thesis, we study the intersection of optimization and machine learning, especially how to use machine learning and optimization tools to make decisions. In Chapter 1, we propose a novel approach for accurate policy evaluation in personalized pricing. We solve an optimization problem to evaluate new pricing strategies, while searching over some worst case revenue functions. In Chapter 2, we consider problems where parameters are predicted using a machine learning model to be used for downstream optimization tasks. Recent works have proposed an integrated approach, accounting for how predictions are used in the downstream optimization problem, instead of just minimizing prediction error. We analyze the asymptotic performance of methods under the integrated and traditional approaches, in the sense of first-order stochastic. We argue that when the model class is rich enough to cover the ground truth, the traditional predict-then-optimize approach outperforms the integrated approach, and the performance ordering between the two approaches is reversed when the model is misspecified. In Chapter 3, we present a new class of architectures for reinforcement learning, Implicit Two-Tower (ITT) policies, where the actions are chosen based on the attention scores of their learnable latent representations with those of the input states. We show that ITT-architectures are particularly suited for evolutionary optimization and the corresponding policy training algorithms outperform their vanilla unstructured implicit counterparts as well as commonly used explicit policies. In  \nChapter 4, we consider an active learning problem, in which the learner has the ability to sequentially select unlabeled samples for labeling. A typical active learning algorithm would sample more points at “difficult\" regions in the feature space to more efficiently use the sampling budget and reduce excess risk. For nonparametric classification with smooth regression functions, we show that nuances in notions of margin that involves the uniqueness of the Bayes classifier, having no apparent effect on rates in passive learning, determine whether or not any active learner can outperform passive learning rates.  \nTable of Contents  \nAcknowledgments ........................................ x  \nIntroduction ........................................... 1  \nChapter 1: Balanced Off-Policy Evaluation for Personalized Pricing ............. 4  \n1.1 Introduction ...................................... 4  \n1.2 Notation and Model .................................. 7  \n1.3 Properties of Weighted Revenue Estimators ..................... 9  \n1.3.1 Mean Squared Error ............................. 9  \n1.3.2 High-Probability Bound ........................... 10  \n1.4 A Balanced Approach for Off-Policy Evaluation in Pricing ............. 11  \n1.4.1 Solution Approach .............................. 13  \n1.5 Theoretical Results .................................. 14  \n1.5.1 Mean Squared Error ............................. 14  \n1.5.2 Bernstein Bound ............................... 15  \n1.6 Numerical Results ................................... 16  \n1.6.1 Mean Squared Error ............................. 16  \n1.6.2 Bernstein Bounds ............................... 21  \n1.7 Hyper-parameter Heuristics .............................. 23  \n1.8 Conclusion ...................................... 25  \nChapter 2: Estimate-then-optimize versus Integrated-estimation-optimization: A Stochastic Dominance Pe","cbCaicsn7Bp1uXHZ","https://ap.wps.com/l/cbCaicsn7Bp1uXHZ","pdf",3157762,1,199,"English","en",105,"# Introduction\n## Chapter 1: Balanced Off-Policy Evaluation for Personalized Pricing\n## Chapter 2: Estimate-then-optimize versus Integrated-estimation-optimization: A Stochastic Dominance Perspective\n## Chapter 3: Implicit Two-Tower Policies\n## Chapter 4: Active Learning with Margin-Related Nuances","[{\"question\":\"What problem does Chapter 1 address in personalized pricing?\",\"answer\":\"Chapter 1 introduces a balanced off-policy evaluation approach for personalized pricing, formulating an optimization problem to evaluate new pricing strategies under worst-case revenue functions.\"},{\"question\":\"How does the thesis compare estimate-then-optimize with integrated estimation–optimization?\",\"answer\":\"It analyzes their asymptotic performance under first-order stochastic dominance, concluding that traditional predict-then-optimize can outperform integrated methods when the model class covers the ground truth, while the ordering can reverse under model misspecification.\"},{\"question\":\"What are Implicit Two-Tower (ITT) policies and why are they useful?\",\"answer\":\"ITT policies choose actions based on attention scores derived from learnable latent representations of input states. The thesis argues they suit evolutionary optimization and that their training algorithms outperform both unstructured implicit counterparts and common explicit policies.\"}]","Integrating Machine Learning and Optimization for Problems in Contextual Decision-Making and Dynamic Learning - Thesis | PDF",1785806518,501,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"integrating-machine-learning-and-optimization-for-problems-in-contextual-decision-making-and-dynamic-learning-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/integrating-machine-learning-and-optimization-for-problems-in-contextual-decision-making-and-dynamic-learning-thesis/121729/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Chapter 1 address in personalized pricing?","Question",{"text":75,"@type":76},"Chapter 1 introduces a balanced off-policy evaluation approach for personalized pricing, formulating an optimization problem to evaluate new pricing strategies under worst-case revenue functions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis compare estimate-then-optimize with integrated estimation–optimization?",{"text":80,"@type":76},"It analyzes their asymptotic performance under first-order stochastic dominance, concluding that traditional predict-then-optimize can outperform integrated methods when the model class covers the ground truth, while the ordering can reverse under model misspecification.",{"name":82,"@type":73,"acceptedAnswer":83},"What are Implicit Two-Tower (ITT) policies and why are they useful?",{"text":84,"@type":76},"ITT policies choose actions based on attention scores derived from learnable latent representations of input states. The thesis argues they suit evolutionary optimization and that their training algorithms outperform both unstructured implicit counterparts and common explicit policies.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]