[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83385-en":3,"doc-seo-83385-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83385,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Ensemble Diversity Optimization for Subjective Supervision","Subjective NLP tasks often show systematic annotator disagreement, so models must represent uncertainty rather than collapsing it to a single label. Ensemble Diversity Optimization (EDO) is a prediction-space framework that end-to-end learns ensemble weights and effective cardinality, while jointly optimizing calibration and utility. EDO uses Gumbel–Softmax relaxation plus a signed diversity regularizer to preserve or suppress disagreement. On four subjective benchmarks, EDO improves probabilistic calibration, reduces cross-entropy and Brier scores, and keeps competitive F1 with better alignment to annotator distributions.","Ensemble Diversity Optimization for Subjective Supervision  \nXia Cui 1 Ziyi Huang2 N. R. Abeynayake 1  \n1 School of Computing and Mathematics, Manchester Metropolitan University, Manchester, UK.  \n2 School of Computer Science, Hubei University, Wuhan, China.  \narXiv :2607 .08493v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nSubjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it. We introduce Ensemble Diversity Optimization (EDO), a prediction-space framework that jointly optimizes ensemble weights, effective cardinality, and calibration through a unified differentiable objective.  \nEDO learns ensemble composition and size endto-end via Gumbel–Softmax relaxation and incorporates a signed diversity regularizer, tuned on validation data, to steer optimization toward either preserving or suppressing disagreement. This regularization prevents ensemble collapse and enables controlled navigation of the utility–calibration trade-off. The framework integrates a soft F1 surrogate, class-weighted cross-entropy to address imbalance, and reliability-weighted diversity to regulate intra-ensemble variability. Experiments on four subjective text-classification benchmarks (ArMIS, ConvAbuse, HS-Brexit, MD-Agreement) show that EDO substantially improves probabilistic calibration, reducing cross-entropy (40–78% depending on baseline) and lowering Brier scores relative to Soft-CE, Soft-MD, Top-5 Voting, and WEL, while maintaining competitive F1 and better alignment with annotator distributions. These results demonstrate that jointly optimizing ensemble structure with a signed diversity regularizer provides an efficient, model-agnostic approach for modeling human subjectivity in supervised learning.  \n1 INTRODUCTION  \nMany NLP tasks exhibit substantial and systematic annotator disagreement. In domains such as content moderation,  \nhate speech detection, and sentiment analysis, divergent annotations arise from semantic ambiguity, contextual dependence, or variation in annotator expertise rather than annotation error [Snow et al., 2008, Plank, 2022, Uma et al., 2022, Cui et al., 2025] . These settings challenge standard supervised learning assumptions, as the target is not a single latent label but a distribution over plausible human interpretations. Nevertheless, prevailing practice aggregates annotations into a single target, discarding distributional information and inducing overfitting to dominant interpretations [Davani et al., 2022, Liu et al., 2022] .  \nFigure 1: Illustration of annotator disagreement.  \nSoft-label supervision preserves annotator label distributions and provides richer training signals [Uma et al., 2020, Rizzi et al., 2024], but continues to optimize a single predictive model. This setting aligns with partial-label learning (PLL) [Cour et al., 2011], where each instance is associated with a set of candidate labels. Unlike classical PLL, which assumes a single hidden ground truth, subjective tasks often involve genuine multiplicity, where multiple labels are simultaneously valid. As a result, existing approaches provide no explicit mechanism for regulating internal predictive variability or distinguishing systematic subjectivity from annotation noise, particularly under class imbalance or heterogeneous annotator reliability.  \nEnsemble methods offer a principled mechanism for representing multiple plausible hypotheses and have been widely used to capture predictive uncertainty [Fort et al., 2019] . However, existing approaches typically rely on fixed architectures and treat diversity as an emergent property rather than an explicit optimization objective. This contrasts with the unified theory of ensemble diversity [Wood et al., 2023],  \nwhich proves that for losses like cross-entropy, ensemble error decomposes exactly into bias, variance and diversity terms. This establishes diversity as a fundamental component of generalization, not an auxiliary heuristi","cbCaicndrRQ1KnSU","https://ap.wps.com/l/cbCaicndrRQ1KnSU","pdf",578987,2,1,21,"English","en",105,"# Abstract\n# Introduction\n## Annotator disagreement in subjective NLP\n## Limits of existing aggregation and soft-label methods\n## Role of ensemble diversity theory\n## Proposed framework: Ensemble Diversity Optimization (EDO)\n## Experimental results and contributions","[{\"question\":\"What problem does EDO address in subjective NLP tasks?\",\"answer\":\"EO addresses systematic annotator disagreement where the supervision reflects a distribution over plausible interpretations rather than a single latent label. It targets uncertainty representation without collapsing disagreements.\"},{\"question\":\"How does EDO optimize ensemble diversity and calibration together?\",\"answer\":\"EDO jointly learns ensemble weights and effective cardinality using a differentiable objective with Gumbel–Softmax relaxation. It adds a signed, reliability-weighted diversity regularizer and balances calibration and utility through class-weighted cross-entropy and a soft F1 surrogate.\"},{\"question\":\"What improvements does EDO achieve on subjective text-classification benchmarks?\",\"answer\":\"On four benchmarks, EDO substantially improves probabilistic calibration, reducing cross-entropy and lowering Brier scores versus multiple baselines while maintaining competitive F1 and stronger alignment with annotator distributions.\"}]",1784187137,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ensemble-diversity-optimization-for-subjective-supervision","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/ensemble-diversity-optimization-for-subjective-supervision/83385/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does EDO address in subjective NLP tasks?","Question",{"text":75,"@type":76},"EO addresses systematic annotator disagreement where the supervision reflects a distribution over plausible interpretations rather than a single latent label. It targets uncertainty representation without collapsing disagreements.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EDO optimize ensemble diversity and calibration together?",{"text":80,"@type":76},"EDO jointly learns ensemble weights and effective cardinality using a differentiable objective with Gumbel–Softmax relaxation. It adds a signed, reliability-weighted diversity regularizer and balances calibration and utility through class-weighted cross-entropy and a soft F1 surrogate.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does EDO achieve on subjective text-classification benchmarks?",{"text":84,"@type":76},"On four benchmarks, EDO substantially improves probabilistic calibration, reducing cross-entropy and lowering Brier scores versus multiple baselines while maintaining competitive F1 and stronger alignment with annotator distributions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]