[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84751-en":3,"doc-seo-84751-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84751,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Attention Limited Reward Learning","Pairwise human comparisons are a core learning interface for modern AI preference modeling. RLHF-style pipelines often assume Bradley–Terry conditional-logit labels reflect latent reward differences, but this assumption misses cases where evaluators have limited information-processing capacity. The paper introduces an attention-scaled reduced-form model that separates ambiguity from value closeness versus ambiguity from hidden evidence detection difficulty. Limited attention can distort what passive comparisons reveal, make learned rankings misleading, and introduces non-scalar, cyclic signatures in real vote data while response time and gaze carry additional gap information.","arXiv :2607 .04590v 1 [ cs .AI] 6 Jul 2026  \nAttention Limited Reward Learning  \nWenqian Xing∗  \nJuly 7, 2026  \nAbstract  \nPairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley–Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption misses through a reduced-form model motivated by rational inattention, in which each label is generated by a low-capacity evaluation channel. The model separates two forms of ambiguity that standard reward modeling tends to conflate: a comparison may be difficult because the two candidates are genuinely close in value, or because the relevant distinction is hard to detect under limited attention. We show that limited attention can fundamentally distort what pairwise comparisons reveal. In particular, passive comparison data cannot generally distinguish reward, attention, and default tendencies, and heterogeneous attention can make standard Bradley–Terry reward modeling recover misleading rankings. Our analysis shows that learning is governed not by the raw number of labels, but by the amount of attended information each label carries. A case study on human votes over language-model pairs from Chatbot Arena exhibits the predicted signature, a cyclic component of the comparison data that exceeds sampling noise and that no scalar reward can represent; a second case study on perceptual comparisons shows that response times and gaze carry gap information that the labels do not. This perspective suggests that human feedback should be treated not as direct revealed preference, but as an attention-limited measurement process: a weak preference signal may reflect hidden evaluation difficulty rather than genuine indifference.  \n1 Introduction  \nHuman feedback has become a central part of how modern AI systems are trained and evaluated. In reinforcement learning from human feedback (RLHF), preference-based policy optimization, and related alignment pipelines, the basic measurement primitive is often simple. A human is shown two candidate responses, rankings, plans, allocations, or trajectories and asked which one is better. A reward model is then trained so that its pairwise differences predict these choices. This template underlies much of contemporary reward learning for language models and agentic systems [7, 23 , 31 , 34 , 36] . Direct preference optimization changes the downstream policy update, but it still relies on the same basic comparison signal linking human choices to reward differences [24] .  \nThe standard statistical abstraction is the Bradley–Terry, or conditional-logit, model [3, 18 , 20] . For a query consisting of a context x and two candidates y 0 , y 1 , it assumes  \nP [y1 ≻ y0 | x] = σ 􀀀r(x, y1 ) − r(x, y0 )􀀁, σ(t) = 1~~ ~~+1e−t . (1)  \nUnder this model, comparison labels are noisy but direct measurements of a single latent reward. If two candidates are chosen at roughly equal rates, the model interprets this as evidence that their  \n∗ Management Science and Engineering Department, Stanford University, [wxing@stanford.edu](wxing@stanford.edu)  \nrewards are nearly equal. Much of the statistical ranking literature studies estimation and ranking in this correctly specified regime [22, 26] .  \nFor alignment, however, the most important comparisons are often not the easiest ones. A response may be fluent but subtly misleading. A plan may appear helpful while creating long-run risks. A model output may satisfy the literal request while violating the user’s underlying intent. In such cases, the relevant distinction is not necessarily small; it may simply be hard to notice. This observation is the premise of work on scalable oversight, which anticipates that as systems become more capable, the outputs whose evaluation matters most are the ones that strain unaided human judgment [1, 2 , 13 , ","cbCaieLGuuAU7dWp","https://ap.wps.com/l/cbCaieLGuuAU7dWp","pdf",659364,1,24,"English","en",105,"# Abstract\n# Introduction\n## Human feedback and comparison-based reward learning\n## Bradley–Terry conditional-logit modeling\n## Limited attention as a structured measurement problem\n## Attention-scaled reduced form and rational inattention","[{\"question\":\"What common modeling assumption do RLHF reward-learning pipelines make about pairwise comparisons?\",\"answer\":\"They often use the Bradley–Terry conditional-logit abstraction, treating choice probabilities as governed by latent reward differences between candidates.\"},{\"question\":\"How does the paper distinguish two types of “hard” comparisons?\",\"answer\":\"A comparison can be hard because candidates are genuinely close in value, or because the reward-relevant distinction is hard to detect under limited attention.\"},{\"question\":\"Why can standard Bradley–Terry reward modeling yield misleading rankings under limited attention?\",\"answer\":\"Because attention differences can distort the information contained in labels, passive comparison data cannot separate reward from attention and default tendencies, and the inferred scalar reward can fail to represent the observed comparison structure.\"}]",1784198034,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"attention-limited-reward-learning","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/attention-limited-reward-learning/84751/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What common modeling assumption do RLHF reward-learning pipelines make about pairwise comparisons?","Question",{"text":74,"@type":75},"They often use the Bradley–Terry conditional-logit abstraction, treating choice probabilities as governed by latent reward differences between candidates.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the paper distinguish two types of “hard” comparisons?",{"text":79,"@type":75},"A comparison can be hard because candidates are genuinely close in value, or because the reward-relevant distinction is hard to detect under limited attention.",{"name":81,"@type":72,"acceptedAnswer":82},"Why can standard Bradley–Terry reward modeling yield misleading rankings under limited attention?",{"text":83,"@type":75},"Because attention differences can distort the information contained in labels, passive comparison data cannot separate reward from attention and default tendencies, and the inferred scalar reward can fail to represent the observed comparison structure.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":28,"slug":107},5,"Comic","comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]