[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127866-en":3,"doc-seo-127866-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127866,2336474466712,"Maeve","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Flexible Machine Learning and Reinforcement Learning in Decision Making","Machine learning and reinforcement learning support decision making, especially in individualized settings where treatment effects vary across subgroups. This dissertation develops flexible methods for estimating optimal Individualized Treatment Rules (ITR) with many treatments, including group-structured outcome weighted learning and an adaptive fusion approach that clusters treatments by similar effects. It further applies reinforcement learning to efficient variable selection in big data via a Markov decision process formulation with dynamic policies and online policy iteration, achieving accurate selection with efficient computation.","FLEXIBLE MACHINE LEARNING AND REINFORCEMENT LEARNING IN  \nDECISION MAKING  \nHaixu Ma  \nA dissertation submitted to the faculty of the University of North Carolina at Chapel Hill in partial ful􀀌llment of the requirements for the degree of Doctor of Philosophy in the Department of Statistics and Operations Research.  \nChapel Hill  \n2024  \nApproved by:  \nYufeng Liu  \nDonglin Zeng  \nGuanting Chen  \nJan Hannig  \nQuoc Tran-Dinh  \n􀀍c2024  \nHaixu Ma  \nALL RIGHTS RESERVED  \nii  \nABSTRACT  \nHaixu Ma: Flexible Machine Learning and Reinforcement Learning in Decision Making  \n(Under the direction of Yufeng Liu and Donglin Zeng)  \nMachine learning and Reinforcement Learning (RL) have received a lot of attentions in decision making problems. For individualized decision making, due to possible signi􀀌cant heterogeneity of treatment e􀀋ects among individuals, decision makers aim to precisely tailor the treatment decision rules to di􀀋erent subgroups of individuals. In the 􀀌rst part of this dissertation, we propose several new approaches to estimate the optimal Individualized Treatment Rules (ITR) when there are many treatments available. For the 􀀌rst project, we introduce the group-structured ITR and propose GRoup Outcome Weighted Learning (GROWL) to estimate the latent structure in the treatment space and the optimal group-structured ITRs through a single optimization. Fisher consistency, the excess risk bound, and the convergence rate of the value function are established to provide a theoretical guarantee for GROWL. For the second project, we propose a novel adaptive fusion based method to cluster the treatments with similar treatment e􀀋ects together and estimate the optimal ITR simultaneously with a single convex optimization. The problem is formulated as balancing loss 􀀀 penalty terms with a tuning parameter, which allows the entire solution path of the treatment clustering process to be clearly visualized hierarchically. Recently, RL has demonstrated its ability to enhance the e􀀎ciency of fundamental problems requiring extensive computational resources. In the second part of this dissertation, we focus on using RL for e􀀎cient variable selection in big data. For the third project, we propose a novel approach, REinforcement learning for Variable Selection (REVS), within the Markov decision process framework. By prioritizing the long-term variable selection accuracy, we propose a dynamic policy to adjust the candidate important variable set, guiding it toward convergence to the true variable set. To enhance computational e􀀎ciency, we present an online policy iteration algorithm integrated with temporal di􀀋erence learning for sequential policy improvement. The REVS is shown to have accurate variable selection with highly e􀀎cient computation.  \nTo my parents, Xiangyang Ma and Ying Sun  \nACKNOWLEDGEMENTS  \nI am immensely grateful for the guidance and mentorship provided by my Ph.D. advisors, Dr. Yufeng Liu and Dr. Donglin Zeng. Working under their mentorship has been my great honor, one that has profoundly shaped my academic journey and research pursuits. Dr. Liu's unwavering support throughout my Ph.D. journey has been instrumental in my career development. He has guided me through the process of modeling complex problems, by starting with intuitive examples before advancing to more general cases. Beyond the technical knowledge, Dr. Liu has imparted invaluable lessons on many soft skills and suggestions for career development. His emphasis on clear communications, meticulous attention to details, and rigorous academic discipline has helped [me navigate the landscape of my Ph.D. career. Equally in](me navigate the landscape of my Ph.D. career. Equally in)􀀍uential, Dr. Zeng's depth of knowledge and hands-on assistance with research ideas and technical discussions have been pivotal in my exploration of challenging research projects. His passion for research is inspiring, motivating me to delve deeper into intriguing research projects with enthusias","cbCaikx1cBIjnqAM","https://ap.wps.com/l/cbCaikx1cBIjnqAM","pdf",1945675,2,1,160,"English","en",105,"# Abstract\n# Acknowledgements\n# Table of Contents\n## List of Tables\n## List of Figures","[{\"question\":\"What is the dissertation’s main focus in decision making?\",\"answer\":\"It studies flexible machine learning and reinforcement learning methods to support decision making, particularly individualized treatment decisions under treatment effect heterogeneity.\"},{\"question\":\"How does the first part estimate optimal Individualized Treatment Rules (ITR)?\",\"answer\":\"It introduces group-structured ITR learning using GROWL to estimate latent structure and optimal group-structured ITRs, and an adaptive fusion method that clusters treatments with similar effects while estimating ITRs via a single convex optimization.\"},{\"question\":\"What is REVS and how does it perform variable selection?\",\"answer\":\"REVS formulates variable selection as a Markov decision process, using a dynamic policy to improve long-term selection accuracy and an online policy iteration algorithm with temporal difference learning to guide sequential improvement efficiently.\"}]","Flexible Machine Learning and Reinforcement Learning in Decision Making | PDF",1785942426,403,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"flexible-machine-learning-and-reinforcement-learning-in-decision-making","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/flexible-machine-learning-and-reinforcement-learning-in-decision-making/127866/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the dissertation’s main focus in decision making?","Question",{"text":76,"@type":77},"It studies flexible machine learning and reinforcement learning methods to support decision making, particularly individualized treatment decisions under treatment effect heterogeneity.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the first part estimate optimal Individualized Treatment Rules (ITR)?",{"text":81,"@type":77},"It introduces group-structured ITR learning using GROWL to estimate latent structure and optimal group-structured ITRs, and an adaptive fusion method that clusters treatments with similar effects while estimating ITRs via a single convex optimization.",{"name":83,"@type":74,"acceptedAnswer":84},"What is REVS and how does it perform variable selection?",{"text":85,"@type":77},"REVS formulates variable selection as a Markov decision process, using a dynamic policy to improve long-term selection accuracy and an online policy iteration algorithm with temporal difference learning to guide sequential improvement efficiently.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]