[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86233-en":3,"doc-seo-86233-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86233,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability","Reinforcement learning from human pairwise trajectory comparisons is extended by allowing experts to label pairs as incomparable, meaning neither trajectory dominates the other. The work formulates a generalized preference-based RL learning problem with explicit solution desiderata. A Bradley–Terry-inspired rationality model is proposed to represent incomparabilities and to infer a multi-dimensional reward function, followed by model property studies. Parameter learning receives a sample complexity analysis under available datasets, and evaluations test reward reconstruction, Pareto frontier recovery, and robustness to varying expert rationality levels.","Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability  \nSimone Drago  \nPolitecnico di Milano, Milan, Italy [simone.drago@polimi.it](simone.drago@polimi.it)  \nMarco Mussi  \nPolitecnico di Milano, Milan, Italy [marco.mussi@polimi.it](marco.mussi@polimi.it)  \nLeonardo Bianconi  \nPolitecnico di Milano, Milan, Italy [leonardo.bianconi@mail.polimi.it](leonardo.bianconi@mail.polimi.it)  \nAlberto Maria Metelli  \nPolitecnico di Milano, Milan, Italy [albertomaria.metelli@polimi.it](albertomaria.metelli@polimi.it)  \narXiv :2607 . 1 1432v 1 [ cs .LG] 13 Jul 2026  \nAbstract  \nIn this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preferencebased RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither trajectory “dominates” the other. We introduce the learning problem and the desiderata that its solution should satisfy.  \nThen, we propose a novel Bradley–Terry-inspired rationality model that effectively captures incomparabilities and infers a multi-dimensional reward function, and we study its properties. We provide a sample complexity analysis for learning the model parameters when a dataset is available. Finally, we evaluate our model’s ability to reconstruct a reward function that aligns with the expert’s comparisons in simulated environments and to recover the Pareto frontier of policies, along with a robustness analysis across varying levels of expert rationality.  \n1 Introduction  \nIn reinforcement learning (RL, Sutton and Barto, 2018), a learning agent interacts with an environment observing a state and, based on it, selecting an action, to optimize a certain objective function over a given time horizon. Crucially, the learning process is guided by a reward function, i.e., a numerical signal provided as feedback to the agent after each action (Sutton, 2004) . With such a formulation, the agent’s goal is to maximize the expected cumulative reward. The reward is often referred to as“the most succinct description of a task”(Ng and Russell, 2000) . However, defining a reward function is a complex endeavor. The system engineer, who designs it, must select numerical values to induce a behavior that “solves” the task, while avoiding unwanted, possibly dangerous, phenomena (e.g., reward hacking, Amodei et al., 2016) . This process, namely reward engineering (Dewey, 2014), is atrial-and-error iterative process that relies on domain knowledge and requires tweaking the reward function, as the induced behavior is highly sensitive to misspecified rewards (Pan et al., 2022) . Preference-based RL (PbRL, Wirth et al., 2017) has emerged as a powerful paradigm for overcoming the inherent reward design challenges of RL. It avoids the reward engineering process entirely, relying instead on pairwise comparisons provided by a human expert and learning the policy that best aligns with them. PbRL has achieved remarkable results in domains ranging from robotics (Lee et al., 2021b) to large language models (Ouyang et al., 2022) . One common approach, namely reinforcement learning from human feedback (RLHF, Christiano et al., 2017) assumes the existence of an underlying, unknown reward (Friedman and Savage, 1952), and comprises two steps: (i) reward model estimation from the expert’s preferences and (ii) policy optimization via traditional RL methods. PbRL’s achievements are largely built on the use of rationality models, which link the underlying (unknown)  \nreward function to observed preferences and guide the stochastic preference generation process. The Preprint.  \nBradley-Terry (BT, Bradley and Terry, 1952) model is most commonly employed in the literature (see, e.g., Christiano et al., 2017; Rafailov et al., 2023; Munos et al., 2024), and it is fed by utility functions that quantify the “goodness” of the alternatives, often represented by the cumulative reward. Und","cbCaicBG2AdBAdUl","https://ap.wps.com/l/cbCaicBG2AdBAdUl","pdf",650280,3,1,28,"English","en",105,"# Introduction\n## Preference-based RL and rationality models\n## Incomparability and multi-dimensional rewards\n## Learning problem and evaluation","[{\"question\":\"How does the paper extend preference-based reinforcement learning beyond standard pairwise comparisons?\",\"answer\":\"It generalizes preference-based RL by introducing a feedback type where the expert can label trajectory pairs as incomparable, i.e., neither option dominates the other.\"},{\"question\":\"What is the proposed rationality model and what does it capture?\",\"answer\":\"The paper introduces a Bradley–Terry-inspired rationality model that captures incomparabilities and infers a multi-dimensional reward function.\"},{\"question\":\"What analyses and experiments are used to validate the approach?\",\"answer\":\"It provides a sample complexity analysis for learning model parameters, then evaluates reward reconstruction accuracy and Pareto frontier recovery in simulated environments, including robustness across different levels of expert rationality.\"}]",1784209686,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"generalizing-preference-based-reinforcement-learning-a-rationality-model-for-incomparability","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/generalizing-preference-based-reinforcement-learning-a-rationality-model-for-incomparability/86233/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the paper extend preference-based reinforcement learning beyond standard pairwise comparisons?","Question",{"text":75,"@type":76},"It generalizes preference-based RL by introducing a feedback type where the expert can label trajectory pairs as incomparable, i.e., neither option dominates the other.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed rationality model and what does it capture?",{"text":80,"@type":76},"The paper introduces a Bradley–Terry-inspired rationality model that captures incomparabilities and infers a multi-dimensional reward function.",{"name":82,"@type":73,"acceptedAnswer":83},"What analyses and experiments are used to validate the approach?",{"text":84,"@type":76},"It provides a sample complexity analysis for learning model parameters, then evaluates reward reconstruction accuracy and Pareto frontier recovery in simulated environments, including robustness across different levels of expert rationality.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]