[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84287-en":3,"doc-seo-84287-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84287,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration","Research on stereotypes in large language models (LLMs) has largely centered on English, leaving non-English and underrepresented cultures under-studied due to missing datasets and costly manual annotation. The paper proposes a cost-efficient human–LLM collaborative annotation framework to construct EspanStereo, a Spanish stereotype dataset covering multiple Spanish-speaking countries. It integrates LLM-generated candidate stereotypes with in-culture annotator validation to capture both established and culturally specific biases. Evaluations show significant cross-country variation, motivating culturally grounded, multilingual assessment beyond English-centric benchmarks.","Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration  \nWeicheng Ma 1* , John Guerrerio2* , and Soroush Vosoughi3  \n1 College of Computing, Georgia Institute of Technology  \n2,3 Computer Science Department, Dartmouth College  \n[1](1 wma76@gatech.edu)[ wma76@gatech.edu](1 wma76@gatech.edu)  \n[2](2 john.j.guerrerio.26@dartmouth.edu)[ john.j.guerrerio.26@dartmouth.edu](2 john.j.guerrerio.26@dartmouth.edu)  \n[3](3 soroush.vosoughi@dartmouth.edu)[ soroush.vosoughi@dartmouth.edu](3 soroush.vosoughi@dartmouth.edu)  \narXiv :2607 .07895v 1 [ cs .CL] 8 Jul 2026  \nAbstract  \nWarning: This paper contains examples of potentially offensive content.  \nResearch on stereotypes in large language models (LLMs) has largely focused on Englishspeaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures. To address this gap, we introduce a cost-efficient human-LLM collaborative annotation framework and apply it to construct EspanStereo, a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries across Europe and Latin America. EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources. Using LLMs to generate candidate stereotypes and in-culture annotators to validate them, we demonstrate the framework’s effectiveness in identifying nuanced, region-specific biases. Our evaluation of Spanish-supporting LLMs using EspanStereo reveals significant variation in stereotypical behavior across countries, highlighting the need for more culturally grounded assessments. Beyond Spanish, our framework is adaptable to other languages and regions, offering a scalable path toward multilingual stereotype benchmarks. This work broadens the scope of stereotype analysis in LLMs and lays the groundwork for comprehensive cross-cultural bias evaluation.  \n1 Introduction  \nThe rise of large language models (LLMs) has advanced computational linguistics but also introduced challenges due to embedded stereotypes. Existing approaches for detecting and mitigating these biases rely on carefully annotated datasets like StereoSet (Nadeem et al., 2021) and CrowS-Pairs (Nangia et al., 2020), which are only in English and reflect stereotypes from a few English-speaking  \n*Equal contribution.  \ncountries, primarily the US. This narrow scope limits research on stereotypes in non-English, often low-resource, cultures. Moreover, stereotypes vary even within the same language. For example, while both the US and the UK are primarily Englishspeaking countries, the stereotype that rural areas are obsessed with guns is US-specific, whereas soccer fanaticism is more associated with the UK. Existing datasets, especially translation-based ones, often overlook such cultural distinctions.  \nComprehensive and culturally diverse stereotype examination datasets are essential to advance stereotype research in LLMs. However, manual data collection, the predominant method for constructing existing datasets (Nadeem et al., 2021 ; Nangia et al., 2020 ; Felkner et al., 2023 ; Zhao et al., 2018), is expensive and labor-intensive, particularly in regions with smaller populations. The most resource-intensive phase of manual data collection is stereotype acquisition, as ensuring country-specific representation requires sufficiently large and diverse participant samples. Constructing country-specific datasets is especially challenging because they rely on a narrower participant pool than datasets spanning an entire language.  \nTo address this challenge, we propose a humanLLM collaborative stereotype annotation framework, which acquires trial stereotypes from LLMs instead of via human annotations. These generated stereotypes are subsequently validated and instantiated by in-culture annotators to ensure quality and accuracy. Using this framework, we construct EspanStereo, a Spanish-language stere","cbCaifJ358WaSy48","https://ap.wps.com/l/cbCaifJ358WaSy48","pdf",8174344,4,1,29,"English","en",105,"# Abstract\n# Introduction\n# Background","[{\"question\":\"Why are existing stereotype datasets insufficient for non-English cultures?\",\"answer\":\"Existing datasets focus on English and a limited set of English-speaking contexts, missing cultural distinctions and often overlooking within-language regional differences. They also require expensive manual collection, making country-specific coverage difficult in smaller or underrepresented cultures.\"},{\"question\":\"How does the proposed human–LLM collaborative framework construct the EspanStereo dataset?\",\"answer\":\"The method acquires trial stereotypes from LLMs instead of relying on human annotation upfront. In-culture annotators then validate and instantiate these candidates to ensure quality and accuracy, resulting in a dataset with country-specific stereotypes.\"},{\"question\":\"What does EspanStereo evaluation reveal about Spanish-supporting LLMs?\",\"answer\":\"Probing and pruning using EspanStereo shows substantial variation across the studied countries in both stereotype prevalence and encoding behaviors. This indicates that regional stereotypes differ and that assessments need country-level cultural grounding.\"}]",1784194596,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scalable-and-culturally-specific-stereotype-dataset-construction-via-human-llm-collaboration","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/scalable-and-culturally-specific-stereotype-dataset-construction-via-human-llm-collaboration/84287/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are existing stereotype datasets insufficient for non-English cultures?","Question",{"text":75,"@type":76},"Existing datasets focus on English and a limited set of English-speaking contexts, missing cultural distinctions and often overlooking within-language regional differences. They also require expensive manual collection, making country-specific coverage difficult in smaller or underrepresented cultures.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed human–LLM collaborative framework construct the EspanStereo dataset?",{"text":80,"@type":76},"The method acquires trial stereotypes from LLMs instead of relying on human annotation upfront. In-culture annotators then validate and instantiate these candidates to ensure quality and accuracy, resulting in a dataset with country-specific stereotypes.",{"name":82,"@type":73,"acceptedAnswer":83},"What does EspanStereo evaluation reveal about Spanish-supporting LLMs?",{"text":84,"@type":76},"Probing and pruning using EspanStereo shows substantial variation across the studied countries in both stereotype prevalence and encoding behaviors. This indicates that regional stereotypes differ and that assessments need country-level cultural grounding.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]