[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86158-en":3,"doc-seo-86158-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86158,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Rank Conditioned Sample Reuse for the Plackett-Luce Best of K Objective","Study a coupled reinforcement-learning objective JWORK = E_{S ~ PL-WORK}[max_{i∈S} R_i], the expected best-of-K reward produced by a size-K Plackett–Luce draw without replacement, under Gumbel-Top-K / Stochastic Beam Search decoding. The paper distinguishes it from the conventional i.i.d. best-of-K target JiidK = E[max_{i≤K} R_i] and shows i.i.d. sample-reuse weights induce bias under the coupled sampler. It builds a rank-conditioned Horvitz–Thompson estimator that reuses a single Gumbel-Top-n pool (n>K) to form unbiased JWORK gradients, with an exact score-function surrogate and a reward-sorted dynamic program reducing an n choose K subset sum to a one-dimensional integral.","arXiv :2607 . 1 1 146v 1 [ cs .LG] 13 Jul 2026  \nRank-Conditioned Sample Reuse for the Plackett–Luce Best-of-K Objective  \nMelveena Jolly Independent Researcher [melveenajollyk@gmail. com](melveenajollyk@gmail. com)  \nMidhun Xavier Independent Researcher [midhunxavier@outlook. com](midhunxavier@outlook. com)  \nJuly 2026  \nAbstract  \nWe study the coupled objective JWORK = ES ∼PL-WORK 􀀂maxi∈S Ri􀀃 : the expected maximum reward of a size-K Plackett–Luce draw without replacement, the law of Gumbel-Top-K / Stochastic Beam Search decoding. This estimand differs from the conventional i.i.d. objective JiidK = E[maxi≤K Ri] (Ri independent) targeted by existing sample-reuse Max@K estimators, and reusing their i.i.d. weights under the coupled sampler is provably biased: we exhibit acsulosamnbedpiales-freordims tfothrheJenO;emyw-rehinwatsatriadtnslcepacwecksiitalishcEassaem[dfle]JOu∇) . θJGOOececntxjoracinibttult-ysicon(poraeisssRt@EoKINinuFstndORanterCiathEteiesstcaaolrnupeadaledrdydrank-conditioned Horvitz–Thompson estimation for the JWORK subset total: from one GumbelTop-n pool (n > K) and its observed priority threshold we build an estimator that reuses all 􀀀 nK􀀁 embedded K-subsets, unbiased for JWORK (Theorem 1) with an unbiased exact score-function surrogate gradient (Proposition 1), together with a reward-sorted Max-specific dynamic program that collapses the 􀀀 nK􀀁-term subset sum (each term carrying a K!-cost set probability)  \nexactly to a one-dimensional integral. A fixed-Q quadrature evaluation costs O (nlog n + nKQ) arithmetic operations (Theorem 2); it is numerically, not algebraically, exact, and we certify no ϵ-approximation rate. Each nonzero degree-K Horvitz–Thompson term has finite second moment exactly when n ≥ 2K, and under the same finite-support interior assumptions the full surrogate gradient has finite second moment whenever n ≥ 2K (Proposition 2); sharpness at the gradient level remains open. The construction recovers the classical single-item priority-sampling estimator at K = 1 . All quantities require only the values and differentiable computation graphs of the n+1 drawn items’ probabilities, so finite structured sequence policies sampled by exact Stochastic Beam Search are covered (Corollary 2) . A certified finite-Q quadrature error bound and countably infinite support remain open.  \n1 Introduction  \nReinforcement learning increasingly optimizes not the mean reward but a best-of-K reward, the quantity a deployed system actually realizes when it samples K candidates and keeps the winner. Two best-of-K estimands must be separated from the first line. The i.i.d. objective,  \nJiidK = E 􀀂max Ri􀀃, R1 ,..., RK i.i.d. from pθ ,  \ni≤K  \nis the standard Max@K / pass@K target and now an active line of work: PKPO (Walder and Karkhanis, 2025), RSPO (Zhang et al., 2025), MaxPO (Takashiro et al., 2026), OrderGrad (Parmas  \net al. , 2026), and related methods derive unbiased policy gradients for it by reusing an i.i.d. batch. This note is about the coupled objective  \nJWORK = ES ∼PL-WORK h axS Rii,  \nthe expected maximum of a size-K Plackett–Luce draw without replacement: the estimand realized when a neural combinatorial optimization (NCO) solver decodes K distinct tours in one GumbelTop-K / Stochastic Beam Search draw and returns the shortest, or an LLM system samples K distinct generations in one such draw and keeps the one a verifier prefers (pass@K under the coupled sampler is the binary-reward special case) . The two estimands differ, and estimators tuned to one are generally biased for the other.  \nThe gap. The cited sample-reuse estimators assume the K samples are drawn i.i.d.: independently, with replacement. Duplicate candidates contribute nothing to the max, so one important design choice is to bias the sampler toward distinct candidates on purpose: Gumbel-Top-k / Stochastic Beam Search (Kool et al. , 2019), deduplication before verification, diverse (penaltybased or deterministic) beam variants, and QMC-coupled bat","cbCaiv62JZ29p9QZ","https://ap.wps.com/l/cbCaiv62JZ29p9QZ","pdf",621080,3,1,26,"English","en",105,"# Abstract\n# Introduction\n## Best-of-K objectives and estimation gap\n## Clarification of sampler-level without replacement","[{\"question\":\"How does JWORK differ from the standard i.i.d. best-of-K objective JiidK?\",\"answer\":\"JWORK is the expected maximum reward from a size-K Plackett–Luce draw without replacement under a coupled sampler, while JiidK assumes K i.i.d. rewards. Estimators designed for the i.i.d. case are generally biased for JWORK under this coupling.\"},{\"question\":\"Why do i.i.d. sample-reuse weights become biased under the Plackett–Luce without-replacement sampler?\",\"answer\":\"Under without-replacement coupling, the joint inclusion probabilities of K-subsets are not the product of marginals. Naively reweighting i.i.d. estimators with Horvitz–Thompson style factors fails, and the paper shows rank-conditioning is required to restore unbiasedness.\"},{\"question\":\"What estimator does the paper propose for JWORK and what key property is guaranteed?\",\"answer\":\"The paper constructs a rank-conditioned Horvitz–Thompson estimator using one Gumbel-Top-n pool (n\\u003eK) and observed priority thresholds, reusing all embedded K-subsets. It guarantees unbiasedness for JWORK (Theorem 1) and provides an unbiased exact score-function surrogate gradient (Proposition 1).\"}]",1784208978,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rank-conditioned-sample-reuse-for-the-plackett-luce-best-of-k-objective","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rank-conditioned-sample-reuse-for-the-plackett-luce-best-of-k-objective/86158/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does JWORK differ from the standard i.i.d. best-of-K objective JiidK?","Question",{"text":75,"@type":76},"JWORK is the expected maximum reward from a size-K Plackett–Luce draw without replacement under a coupled sampler, while JiidK assumes K i.i.d. rewards. Estimators designed for the i.i.d. case are generally biased for JWORK under this coupling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do i.i.d. sample-reuse weights become biased under the Plackett–Luce without-replacement sampler?",{"text":80,"@type":76},"Under without-replacement coupling, the joint inclusion probabilities of K-subsets are not the product of marginals. Naively reweighting i.i.d. estimators with Horvitz–Thompson style factors fails, and the paper shows rank-conditioning is required to restore unbiasedness.",{"name":82,"@type":73,"acceptedAnswer":83},"What estimator does the paper propose for JWORK and what key property is guaranteed?",{"text":84,"@type":76},"The paper constructs a rank-conditioned Horvitz–Thompson estimator using one Gumbel-Top-n pool (n>K) and observed priority thresholds, reusing all embedded K-subsets. It guarantees unbiasedness for JWORK (Theorem 1) and provides an unbiased exact score-function surrogate gradient (Proposition 1).","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]