[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122328-en":3,"doc-seo-122328-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122328,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Learning the Pareto Set Under Incomplete Preferences - Pure Exploration in Vector Bandits","This work studies pure exploration in multi-objective bandit problems with vector-valued rewards, aiming to (approximately) identify the Pareto set of arms despite incomplete preferences modeled by a polyhedral convex cone. The paper targets the open challenge of sample-efficient learning algorithms for these settings. It introduces Pareto Vector Bandits (PaVeBa), an adaptive elimination method that nearly attains the gap-dependent and worst-case lower bounds for (ϵ,δ)-PAC Pareto set identification. Numerical experiments further compare PaVeBa and heuristic variants with state-of-the-art multi-objective and vector optimization methods on real-world datasets.","Learning the Pareto Set Under Incomplete Preferences: Pure  \nExploration in Vector Bandits  \nEfe Mert Karag¨ozl¨u  \nBilkent University, Ankara, Turkey  \nC¸a˘gın Ararat  \nBilkent University, Ankara, Turkey  \nYa¸sar Cahit Yıldırım  \nBilkent University, Ankara, Turkey  \nCem Tekin  \nBilkent University, Ankara, Turkey  \nAbstract  \nWe study pure exploration in bandit problems with vector-valued rewards, where the goal is to (approximately) identify the Pareto set of arms given incomplete preferences induced by a polyhedral convex cone. We address the open problem of designing sampleefficient learning algorithms for such problems. We propose Pareto Vector Bandits (PaVeBa), an adaptive elimination algorithm that nearly matches the gap-dependent and worst-case lower bounds on the sample complexity of (ϵ,δ)-PAC Pareto set identification.  \nFinally, we provide an in-depth numerical investigation of PaVeBa and its heuristic variants by comparing them with the state-ofthe-art multi-objective and vector optimization algorithms on several real-world datasets with conflicting objectives.  \n1 INTRODUCTION  \nPure exploration seeks to identify the optimal arms through sequential interaction. Usually, the error upper bound for identification is given and one seeks to minimize the sampling budget. This approach is termed the fixed confidence setting (Karnin et al. , 2013) . Notable algorithms for this include Exponential Gap Elimination (Karnin et al., 2013) and Track-andStop strategy (Garivier and Kaufmann, 2016) . A similar concept is the (ϵ,δ)-probably approximately correct (PAC) best arm identification introduced by Even-Dar  \nProceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS) 2024, Valencia, Spain. PMLR: Volume 238 . Copyright 2024 by the author(s) .  \net al. (2006), where the aim is to find an ϵ-optimal arm with 1 − δ confidence. At ϵ = 0, it coincides with the fixed confidence setting.  \nAlthough traditional pure exploration approaches primarily focus on scalar rewards, many real-world exploration challenges present multiple competing objectives. For instance, a communication channel with a low error rate, high bit rate, and a narrow bandwidth tends to consume more power; and a more complex neural network tends to require more time, computation, and data to be trained. Therefore, singleobjective approaches fall short for real problems with D > 1 objectives that cannot be simplified to scalar optimization. Such problems gained attention from bandit literature in applications like digital hardware design (Zuluaga et al., 2016) and treatment optimization (Lizotte and Laber, 2016) . To handle these, one needs a vector-valued multi-armed bandit framework that extends scalar rewards. Although vector rewards can be scalarized by a weighted sum of individual objectives, this approach can be difficult for the practitioner since it requires choosing a weight vector. Furthermore, using weighted linear combinations is not the only way to scalarize the reward vectors, and each real-life problem may require its own wise choice of (possibly highly nonlinear) scalarization function; hence, scalarization becomes even harder, and it motivates the study of exploration among vector-valued bandits on its own.  \nTo that end, some work on bandit literature focused on the identification of the Pareto set of arms, i.e., arms that are not dominated by any other arm in all objectives. Noteworthy contributions include algorithms for (ϵ,δ)-PAC Pareto set identification (Auer et al., 2016) defined similarly to the single-objective case, exploration using Gaussian processes in large datasets (Zuluaga et al., 2016; Shah and Ghahramani, 2016), feasible arm identification (Katz-Samuels and Scott, 2018),  \nlog OPC  \nFigure 1: Three anesthetics ordered by a large polyhedral cone. The only Pareto optimal drug is Diazepam.(Data gathered from Wishart et al. (2018) and Huang et al. (2022).)  \nand algorithms using information-theor","cbCaihdMlUEo91Kh","https://ap.wps.com/l/cbCaihdMlUEo91Kh","pdf",6052997,1,39,"English","en",105,"# Abstract\n# Introduction\n## Fixed confidence vs (ϵ,δ)-PAC best arm identification\n## Vector-valued bandits and scalarization challenges\n## Pareto set identification under standard componentwise order\n## Conic dominance and incomplete preferences\n## Illustrative example in anesthetic design","[{\"question\":\"What is the main goal of the paper in vector bandit pure exploration?\",\"answer\":\"The goal is to approximately identify the Pareto set of arms under incomplete preferences induced by a polyhedral convex cone.\"},{\"question\":\"Why do standard scalar-reward pure exploration methods fall short for real problems?\",\"answer\":\"Real exploration tasks often involve multiple competing objectives that cannot be reduced to scalar optimization without choosing difficult or problem-specific scalarization functions.\"},{\"question\":\"What is PaVeBa and what theoretical result does it target?\",\"answer\":\"PaVeBa is an adaptive elimination algorithm designed to nearly match the known gap-dependent and worst-case lower bounds for (ϵ,δ)-PAC Pareto set identification.\"}]","Learning the Pareto Set Under Incomplete Preferences - Pure Exploration in Vector Bandits | PDF",1785810033,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-the-pareto-set-under-incomplete-preferences-pure-exploration-in-vector-bandits","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-the-pareto-set-under-incomplete-preferences-pure-exploration-in-vector-bandits/122328/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the paper in vector bandit pure exploration?","Question",{"text":75,"@type":76},"The goal is to approximately identify the Pareto set of arms under incomplete preferences induced by a polyhedral convex cone.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do standard scalar-reward pure exploration methods fall short for real problems?",{"text":80,"@type":76},"Real exploration tasks often involve multiple competing objectives that cannot be reduced to scalar optimization without choosing difficult or problem-specific scalarization functions.",{"name":82,"@type":73,"acceptedAnswer":83},"What is PaVeBa and what theoretical result does it target?",{"text":84,"@type":76},"PaVeBa is an adaptive elimination algorithm designed to nearly match the known gap-dependent and worst-case lower bounds for (ϵ,δ)-PAC Pareto set identification.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]