[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123524-en":3,"doc-seo-123524-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123524,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Learning Parametric Distributions from Samples and Preferences - Research paper","Recent advances in language modeling highlight preference feedback as a key signal for improving model performance. This paper characterizes when preference information enables better parameter estimation for continuous parametric distributions. The learner receives sample pairs from an unknown distribution and relative preferences determined by a shared unknown parameter. Preference-based M-estimators reduce asymptotic variance over sample-only M-estimators, with deterministic preferences enabling an O(1/n) estimation error rate. A matching lower bound is provided, with assumptions covering cases such as Gaussian or Laplace rewards.","Learning Parametric Distributions from Samples and Preferences  \nMarc Jourdan 1 Gizem Yüce 1 Nicolas Flammarion 1  \nAbstract  \nRecent advances in language modeling have underscored the role of preference feedback in enhancing model performance. This paper investigates the conditions under which preference feedback improves parameter estimation in classes of continuous parametric distributions. In our framework, the learner observes pairs of samples from an unknown distribution along with their relative preferences depending on the same unknown parameter. We show that preference-based M-estimators achieve a better asymptotic variance than sample-only M-estimators, further improved by deterministic preferences. Leveraging the hard constraints revealed by deterministic preferences, we propose an estimator achieving an estimation error scaling of O(1/n)—a significant improvement over the Θ(1/ √n) rate attainable with samples alone. Next, we establish a lower bound that matches this accelerated rate; up to dimension and problem-dependent constants. While the assumptions underpinning our analysis are restrictive, they are satisfied by notable cases such as Gaussian or Laplace distributions for preferences based on the log-probability reward.  \n1. Introduction  \nRecent progress in language modeling has showcased the effectiveness of preference feedback for fine-tuning (Ziegler et al., 2019 ; Ouyang et al., 2022 ; Bai et al., 2022 ; Touvron et al., 2023 ; Dubey et al., 2024) . Preference data—indicating relative quality between outcomes—consistently outperforms approaches using positive examples only like supervised fine-tuning (Ivison et al., 2024) . This empirical success suggests that preference feedback introduces new, complementary information beyond the observed data. Understanding how and why preferences provide this advantage requires connecting the preference model to the  \n1 School of Computer and Communication Sciences, EPFL, Lausanne, Vaud, Switzerland. Correspondence to: Marc Jourdan \u003C[marc.jourdan@epfl.ch](marc.jourdan@epfl.ch)>.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \ndata-generating process (Ge et al., 2024) .  \nTo understand the role of preference feedback, we focus on a simpler yet illustrative problem: parameter estimation for parametric distributions and preferences. Specifically, the learner observes pairs of samples from an unknown distribution, along with preferences informed by the same parameter. For instance, preferences based on log-probabilities naturally link the preference and probability models, though other formulations are possible (Huang et al., 2024) .  \nFor continuous distributions, we uncover a significant statistical learning gap between preference-based and sampleonly estimators. This paper primarily investigates this gap, taking the sample-only maximum likelihood estimator (MLE)—optimal among unbiased estimators—as a baseline. The well-established theory of M-estimators (Van der Vaart, 2000) suggests that preference-based M-estimators improve asymptotic variance under certain conditions. Yet, this improvement is modest: when samples are of similar quality, preference feedback approaches a fair coin toss, providing minimal additional information. While reducing asymptotic variance is encouraging, it does not fully explain the substantial performance gains observed empirically in large-scale language models.  \nFor deterministic preferences, we prove a more striking result: preference-based estimators achieve a statistically significant acceleration in parameter estimation. Specifically, we show that the estimation error scales as O(1/n) instead of the O(1/ √n) rate achieved by sample-only estimators. This acceleration is supported by a matching lower bound, up to dimension and problem-dependent constants.  \nWhile this acceleration might sound surprising, the Θ(1/n) rate can already be observed ","cbCaimxVyvvbjzg0","https://ap.wps.com/l/cbCaimxVyvvbjzg0","pdf",1718198,1,28,"English","en",105,"# Abstract\n# 1. Introduction\n## 1.1. Contributions","[{\"question\":\"What problem does the paper investigate?\",\"answer\":\"It studies how preference feedback affects parameter estimation for continuous parametric distributions. The learner observes sample pairs and preferences derived from the same unknown parameter.\"},{\"question\":\"How do preference-based estimators compare with sample-only estimators?\",\"answer\":\"Preference-based M-estimators achieve a better asymptotic variance than sample-only M-estimators. With deterministic preferences, they further accelerate estimation error scaling to O(1/n).\"},{\"question\":\"What role do deterministic preferences play in the results?\",\"answer\":\"Deterministic preferences impose hard constraints on feasible parameters. These constraints enable the accelerated O(1/n) estimation error rate, supported by a matching lower bound.\"}]","Learning Parametric Distributions from Samples and Preferences - Research paper | PDF",1785817097,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-parametric-distributions-from-samples-and-preferences-research-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-parametric-distributions-from-samples-and-preferences-research-paper/123524/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper investigate?","Question",{"text":75,"@type":76},"It studies how preference feedback affects parameter estimation for continuous parametric distributions. The learner observes sample pairs and preferences derived from the same unknown parameter.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do preference-based estimators compare with sample-only estimators?",{"text":80,"@type":76},"Preference-based M-estimators achieve a better asymptotic variance than sample-only M-estimators. With deterministic preferences, they further accelerate estimation error scaling to O(1/n).",{"name":82,"@type":73,"acceptedAnswer":83},"What role do deterministic preferences play in the results?",{"text":84,"@type":76},"Deterministic preferences impose hard constraints on feasible parameters. These constraints enable the accelerated O(1/n) estimation error rate, supported by a matching lower bound.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]