[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-203800-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-203800-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","low-frequency-names-exhibit-bias-and-overfitting-in-contextualizing-language-models","Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models","","Examines how training-corpus frequency influences tokenization, contextualization, representational stability, and bias in BERT, GPT-2, T5, and XLNet using a U.S. first-name dataset labeled by predominant gender and racial group. Predominantly female and non-white names appear less often in model training corpora. Infrequent names show higher self-similarity across contexts and lower similarity to initial representations, with frequency–self-similarity and frequency–CKA correlations reaching −0.763 and up to −0.702. Racial bias correlates with lower-frequency minority names in BERT, linking them to unpleasantness.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/low-frequency-names-exhibit-bias-and-overfitting-in-contextualizing-language-models/203800/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/low-frequency-names-exhibit-bias-and-overfitting-in-contextualizing-language-models/203800.png","ImageObject",300,407,{"name":42,"@type":43},"Ethan Miller","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-06","2026-09-04",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",7,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"Which language models are analyzed for name-frequency effects in the study?","Question",{"text":63,"@type":64},"The study evaluates BERT, GPT-2, T5, and XLNet on how first-name frequency affects tokenization, contextualization, similarity behavior, and bias.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does training-corpus frequency relate to self-similarity of infrequent names?",{"text":68,"@type":64},"Infrequent names become more self-similar across contexts; the study reports a Spearman correlation between frequency and self-similarity as low as −0.763.",{"name":70,"@type":61,"acceptedAnswer":71},"What bias pattern does the paper find for minority-group names in BERT?",{"text":72,"@type":64},"Lower-frequency minority names are more associated with unpleasantness in BERT, with a reported Spearman correlation of about −0.492 between racial bias and name frequency.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},203800,1788563977,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,110,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":107,"show_sort_weight":108,"slug":109},"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},8,30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":111,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},687207017582,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","Low Frequency Names Exhibit Bias and Overﬁtting in Contextualizing  \nLanguage Models  \nRobert Wolfe  \nUniversity of Washington [rwolfe3@uw.edu](rwolfe3@uw.edu)  \nAylin Caliskan  \nUniversity of Washington [aylin@uw.edu](aylin@uw.edu)  \nAbstract  \nWe use a dataset of U.S. ﬁrst names with labels based on predominant gender and racial group to examine the effect of training corpus frequency on tokenization, contextualization, similarity to initial representation, and bias in BERT, GPT-2, T5, and XLNet. We show that predominantly female and non-white names are less frequent in the training corpora of these four language models. We ﬁnd that infrequent names are more self-similar across contexts, with Spearman's 􀀚 between frequency and self-similarity as low as 􀀀:763 . Infrequent names are also less similar to initial representation, with Spearman's 􀀚 between frequency and linear centered kernel alignment (CKA) similarity to initial representation as high as :702 . Moreover, we ﬁnd Spearman's  \n􀀚 between racial bias and name frequency in BERT of :492, indicating that lower-frequency minority group names are more associated with unpleasantness. Representations of infrequent names undergo more processing, but are more self-similar, indicating that models rely on less context-informed representations ofuncommon and minority names which are overﬁt to a lower number of observed contexts.  \n1 Introduction  \nHuman social perception is linked to frequency of observation. Hughes et al. (2019) show using functional magnetic resonance imaging (fMRI) scans that humans are more aware of variation in the faces of members of their own race, and perceive members of other races as repeated instances of asocial class, rather than as individuals. Most people interact more with individuals of their own race, and develop better cognitive skills for differentiating members of the race they see most frequently. Recent research indicates that state-of-the-art AI systems mirror such biased and unequal human perceptions. For example, Buolamwini and Gebru (2018) show that under-representation in the training data of computer vision models causes poor  \nperformance on classiﬁcation tasks for women and minority racial groups.  \nFirst names are used in both social psychology and Natural Language Processing (NLP) as a proxy to ground truth data for studying racial and gender biases. The implicit association test (IAT) of Greenwald et al. (1998) ﬁnds that study participants perceive European-American names as more pleasant than African-American names, and Caliskan et al.(2017) demonstrate with the word embedding association test (WEAT) that human biases observed in the IAT exist in static word embeddings, learned vector representations of words widely used in NLP. We use a list of ﬁrst names labeled by gender and racial group based on U.S. Social Security Administration data and the names dataset of Tzioumis (2018) to analyze how name frequency affects minority social groups in four neural language models: BERT, GPT-2, XLNet, and T5 . Neural language models have advanced the state of the art in NLP, and are found in consequential NLP contexts such as Google Search (Nayak, 2019) . These models produce contextualized word embeddings, which incorporate information from the context in which the word occurs over a series of neural network layers. May et al. (2019) and Guo and Caliskan (2021) show that racial, gender, and intersectional biases exist in neural language models. We examine whether under-representation in the training corpora of such models causes them to overﬁt nonwhite and female names, reducing model generalization for underrepresented minorities. We list our research questions and contributions:  \n(1) Are minority social group names less frequent in the training corpora of neural language models? We process four training corpora and ﬁnd that white and male names are the most frequent in all training corpora.  \n(2) Are infrequent and minority group names ","cbCaiuB6HZ5vfVg8","https://ap.wps.com/l/cbCaiuB6HZ5vfVg8","pdf",428357,15,"English","# Abstract\n# Introduction\n## Research questions and contributions","[{\"question\":\"Which language models are analyzed for name-frequency effects in the study?\",\"answer\":\"The study evaluates BERT, GPT-2, T5, and XLNet on how first-name frequency affects tokenization, contextualization, similarity behavior, and bias.\"},{\"question\":\"How does training-corpus frequency relate to self-similarity of infrequent names?\",\"answer\":\"Infrequent names become more self-similar across contexts; the study reports a Spearman correlation between frequency and self-similarity as low as −0.763.\"},{\"question\":\"What bias pattern does the paper find for minority-group names in BERT?\",\"answer\":\"Lower-frequency minority names are more associated with unpleasantness in BERT, with a reported Spearman correlation of about −0.492 between racial bias and name frequency.\"}]","Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models | PDF",38]