[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84292-en":3,"doc-seo-84292-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84292,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","When Debiasing Backfires Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation","Preprocessing-based stereotype mitigation methods in NLP, such as training on debiased corpora or editing datasets before/after learning, often reduce measured bias for targeted groups. Yet this work finds unintended shifts: stereotyping or counter-stereotyping can rise relative to neutral baselines for other demographics, even across unrelated categories. Experiments cover encoder-only and decoder-only models, multiple preprocessing strategies, and varied data scales on Wikipedia.","When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation  \nYahan Zheng  \nDartmouth College [yahan.zheng.gr@dartmouth.edu](yahan.zheng.gr@dartmouth.edu)  \nSoroush Vosoughi  \nDartmouth College [soroush.vosoughi@dartmouth.edu](soroush.vosoughi@dartmouth.edu)  \nJohn J. Guerrerio  \nDartmouth College [john.j.guerrerio.26@dartmouth.edu](john.j.guerrerio.26@dartmouth.edu)  \nWeicheng Ma  \nOakland University [weichengma@oakland.edu](weichengma@oakland.edu)  \narXiv :2607 .07937v 1 [ cs .CL] 8 Jul 2026  \nAbstract  \nPreprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP. While these approaches reduce measurable stereotypes for targeted groups, we find they often induce unintended shifts—side effects, where stereotyping or counter-stereotyping can increase relative to neutral baselines for other demographics, including across unrelated demographic categories. We demonstrate these side effects across two model families (encoderonly and decoder-only), multiple preprocessing strategies (removing stereotypical sentences, removing group mentions, and swapping group references), and both pre-and post-training at different data scales on Wikipedia. Standard benchmarks frequently miss these shifts. Using attention-rollout analysis, we observe that such side effects are not accompanied by large changes in attention flow, complicating mechanistic explanations. We discuss implications for evaluation, provide actionable diagnostics, and argue for side-effect-aware, transparent mitigation practices.  \n1 Introduction  \nPre-trained language models (PLMs) encode and propagate social stereotypes (Bolukbasi et al., 2016 ; Caliskan et al., 2017 ; Zhao et al., 2019), raising concerns about safe and responsible deployment. Among mitigation strategies, preprocessingbased methods seek to modify training data, for instance, by removing or altering stereotypical content, because they are simple to apply and impose no inference-time cost (Gallegos et al., 2024) . Intuitively, such interventions should hypothetically prevent models from learning unwanted associations by isolating them from stereotypical patterns. For consistency, we use the term stereotype to refer to any unwanted social bias encoded by language models from their training data, and mitigation or  \ndebiasing to refer to interventions aimed at reducing these encoded stereotypes.  \nDespite the intuitive appeal and partial effectiveness of preprocessing-based methods, to date, no PLM has achieved complete freedom from stereotypes, raising questions about their true efficacy. More importantly, it is unclear whether datalevel mitigation eliminates harmful associations or merely redistributes them.  \nWe revisit this question empirically by curating debiased Wikipedia corpora targeting six demographic groups spanning three categories (gender, race, religion) . We use these corpora for both preand post-training of two PLMs (encoder-only TinyBERT (Jiao et al., 2020) and decoder-only GPT-2 (Radford et al., 2019)) and then measure changes in stereotype expression toward all groups. We report three findings:  \n1. Unintended shifts within and across categories. While stereotyping toward the group targeted for debiasing decreases, we frequently observe increased stereotyping or counterstereotyping for other groups (including across bias categories) . We define this phenomenon as a side effect: a case where a mitigation method decreases stereotype measures for the target group but simultaneously induces undesired changes for one or more non-target groups. Notably, these shifts are often asymmetric and hard to anticipate. We quantify trends on StereoSet and CrowS-Pairs (Nadeem et al., 2021 ; Nangia et al., 2020) . These patterns cannot be explained solely by changes in the distribution of stereotypical/anti-stereotypical training examples.  \n2. Robustness across settings. Side effects appear under three ","cbCaidfeypIN4xNg","https://ap.wps.com/l/cbCaidfeypIN4xNg","pdf",526077,4,1,21,"English","en",105,"# Introduction\n# Background","[{\"question\":\"What phenomenon do the authors report about preprocessing-based debiasing?\",\"answer\":\"They report unintended side effects: mitigation can reduce stereotyping for the target group while increasing stereotyping or counter-stereotyping for non-target groups, sometimes across different bias categories.\"},{\"question\":\"How do they test side effects across models and training setups?\",\"answer\":\"They curate debiased Wikipedia corpora for six demographic groups and apply both pre-training and post-training to encoder-only TinyBERT and decoder-only GPT-2, using multiple preprocessing strategies and different training data scales.\"},{\"question\":\"Why do the authors argue existing benchmarks and analysis may miss these effects?\",\"answer\":\"Standard benchmarks often fail to surface the shifts, and attention-rollout analysis shows only small changes in attention flow, making attention-based mechanistic explanations insufficient.\"}]",1784194617,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-debiasing-backfires-counterintuitive-side-effects-of-preprocessing-based-stereotype-mitigation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/when-debiasing-backfires-counterintuitive-side-effects-of-preprocessing-based-stereotype-mitigation/84292/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What phenomenon do the authors report about preprocessing-based debiasing?","Question",{"text":75,"@type":76},"They report unintended side effects: mitigation can reduce stereotyping for the target group while increasing stereotyping or counter-stereotyping for non-target groups, sometimes across different bias categories.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do they test side effects across models and training setups?",{"text":80,"@type":76},"They curate debiased Wikipedia corpora for six demographic groups and apply both pre-training and post-training to encoder-only TinyBERT and decoder-only GPT-2, using multiple preprocessing strategies and different training data scales.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do the authors argue existing benchmarks and analysis may miss these effects?",{"text":84,"@type":76},"Standard benchmarks often fail to surface the shifts, and attention-rollout analysis shows only small changes in attention flow, making attention-based mechanistic explanations insufficient.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]