[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85228-en":3,"doc-seo-85228-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85228,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Demographic Prompting at Scale: When More Attributes Hurt LLM–Human Agreement","Investigates how annotator demographic attributes, used as prompt cues, affect alignment between large language model (LLM) predictions and human annotations across five subjective NLP tasks. Five open-source LLMs are evaluated while varying demographic prompt composition from single-attribute to full-attribute settings. Results show alignment peaks with one to three high-signal attributes and declines under over-specification. Influence magnitude alone does not predict gains; learnability and directional coherence matter. Neuron probing links specialized activation to alignment only with coherent annotation signals, indicating steerability is not guaranteed by activation volume.","Demographic Prompting at Scale: When More Attributes Hurt  \nLLM–Human Agreement  \nMahammed Kamruzzaman, Shrabon Kumar Das, Gene Louis Kim  \nBellini College of AI, Cybersecurity and Computing University of South Florida {kamruzzaman1, das157, genekim,}@[usf.edu](usf.edu)  \narXiv :2607 . 10590v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nWe investigate how annotator demographic attributes, supplied as prompt cues, shape the alignment between large language model (LLM) predictions and human annotations across five tasks. Using five open-source LLMs, we systematically vary the number and composition of demographic components in the prompt, spanning every combination from single-attribute through full-attribute configurations. Our experiments reveal three principal findings. First, alignment consistently peaks with one to three high-signal attributes and degrades under the full attribute set, establishing a clear over-specification threshold. Second, the overall magnitude of demographic influence on human annotations does not predict which attributes improve LLM alignment; instead, both the learnability and the directional coherence of each attribute’s annotation signal need to be considered jointly. Third, neuron probing reveals that specialized activation correlates with alignment gains only under coherent annotation signals, and that activation volume alone does not imply steerability. Together, these results demonstrate that demographic prompting is not a monolithic intervention: its utility is highly context-dependent, shaped by attribute signal quality, task characteristics, and model architecture.  \n1 Introduction  \nLLMs are increasingly used as substitutes for, or complements to, human annotators on subjective NLP tasks such as toxicity detection, sentiment analysis, and offensiveness rating (Gilardi et al., 2023 ; Törnberg, 2023) . Because these tasks reflect annotator subjectivity, a natural question arises: how well can an LLM actually model a given demographic perspective when prompted to do so? A growing body of work shows that LLM outputs are not demographically neutral, predictions tend to align more closely with certain demographic  \ngroups in the absence of demographic cues, and explicitly incorporating such cues into the prompt can shift model behaviour in ways that are neither uniform nor always beneficial (Beck et al., 2024 ; Sun et al., 2025 ; Alipour et al., 2025 ; Kamruzzamanet al., 2025 ; Schäfer et al., 2025) .  \nDespite this progress, the existing literature leaves several important questions unresolved. Most prior studies examine only one or two specific demographic attributes at a time, or compare a no-demographic baseline against a single all-attributes prompt, without exploring the combinatorial space in between (Beck et al., 2024 ; Sun et al., 2025 ; Alipour et al., 2025) . None systematically chart the incremental trajectory from singleattribute through multi-attribute to full-attribute prompting, nor investigate the signal to noise relationship along this trajectory. To address this, we pose our first research question: RQ1: To what extent do individual versus combined annotator demographic attributes shape LLM–human alignment, which demographic features (or combinations) most significantly affect alignment, and is there a threshold beyond which additional demographic information no longer benefits alignment?  \nWhile several studies report that demographic prompting sometimes helps and sometimes hurts alignment (Gupta et al., 2023 ; Brown et al., 2025 ; Kamruzzaman et al., 2024), the field lacks a principled account of why: what properties of a demographic attribute, at the dataset level, predict whether prompting with it will improve a given model’s agreement with human annotators? An attribute may strongly predict variation in human labels yet carry internally opposed subgroup signals that no single persona prompt can resolve. Understanding this requires moving beyond aggregate importance measures to cha","cbCailymT9q4IDxs","https://ap.wps.com/l/cbCailymT9q4IDxs","pdf",1673743,3,1,26,"English","en",105,"# Abstract\n# Introduction\n## Research Questions (RQ1–RQ3)\n## Experimental Setup and Contributions","[{\"question\":\"How do annotator demographic attributes in prompts affect LLM–human agreement?\",\"answer\":\"Demographic prompting influences alignment between LLM outputs and human annotations, and the effect changes with the number and composition of attributes. Alignment peaks with one to three high-signal attributes and degrades when the full attribute set is used.\"},{\"question\":\"Why doesn’t the overall strength of demographic influence predict alignment gains?\",\"answer\":\"The document states that alignment improvements depend jointly on how learnable each attribute’s annotation signal is and whether its signal is directionally coherent. The magnitude of influence on human labels alone is insufficient.\"},{\"question\":\"What do neuron probing results suggest about internal mechanisms?\",\"answer\":\"Specialized neuron activation correlates with alignment gains only when the corresponding annotation signals are coherent. High activation volume can occur without resulting in steerability or improved agreement.\"}]",1784201872,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"demographic-prompting-at-scale-when-more-attributes-hurt-llmhuman-agreement","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/demographic-prompting-at-scale-when-more-attributes-hurt-llmhuman-agreement/85228/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How do annotator demographic attributes in prompts affect LLM–human agreement?","Question",{"text":75,"@type":76},"Demographic prompting influences alignment between LLM outputs and human annotations, and the effect changes with the number and composition of attributes. Alignment peaks with one to three high-signal attributes and degrades when the full attribute set is used.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why doesn’t the overall strength of demographic influence predict alignment gains?",{"text":80,"@type":76},"The document states that alignment improvements depend jointly on how learnable each attribute’s annotation signal is and whether its signal is directionally coherent. The magnitude of influence on human labels alone is insufficient.",{"name":82,"@type":73,"acceptedAnswer":83},"What do neuron probing results suggest about internal mechanisms?",{"text":84,"@type":76},"Specialized neuron activation correlates with alignment gains only when the corresponding annotation signals are coherent. High activation volume can occur without resulting in steerability or improved agreement.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]