[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83281-en":3,"doc-seo-83281-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83281,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","DiaLLM: Investigating the Robustness-Generation Gap in English Dialect Adaptation","Large language models can understand dialectal English yet still generate output biased toward standard, US-leaning forms, leaving dialectal generation insufficiently addressed. This work introduces DiaLLM, which continually pretrains three open-weight model families on the International Corpus of English and applies implicit and explicit post-training, each combined with multiple alignment strategies, enabling controlled comparisons across Australian, Indian, and Northern British English.","DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation  \nJordan Painter 1 Dipankar Srirag2 Adarsh Kappiyath 1 Diptesh Kanojia 1 Aditya Joshi2 Lu Yin 1  \n1Institute for People-Centered AI, University of Surrey, Surrey, United Kingdom  \n2University of New South Wales, Sydney, Australia  \n{j.painter,a.kappiyath,d.kanojia, [l.yin}@surrey.ac.uk](l.yin}@surrey.ac.uk)[ ](l.yin}@surrey.ac.uk){d.srirag,[aditya.joshi}@unsw.edu.au](aditya.joshi}@unsw.edu.au)  \narXiv :2607 .07669v 1 [ cs .CL] 8 Jul 2026  \nAbstract  \nLarge language models increasingly understand dialectal English, yet still produce only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce DiaLLM, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal that dialectal robustness and generation are dissociated:  \nbenchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and preferred over broad alignment, yet the method that most aggressively optimises the dialectal reward is not preferred by human evaluators. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.  \n1 Introduction  \nLarge language models (LLMs) achieve strong performance across a range of NLP tasks (Wanget al., 2024 ; Touvron et al., 2023), but often falter on dialectal and non-standard varieties—regional and social forms that diverge from standardised norms (Joshi et al., 2025 ; Blodgett et al., 2016 ; Hofmann et al., 2024 ; Mire et al., 2025) . These disparities reflect structural imbalances in training data: large-scale corpora privilege majority language forms, marginalising non-standard varieties (Bender et al., 2021 ; Joshi et al., 2020) . While domain adaptation has received substantial attention  \nFigure 1: Overview of the DiaLLM pipeline: continual pretraining on ICE followed by either implicit adaptation (standard SFT + alignment) or explicit adaptation (dialectal SFT + variety-targeted alignment), with DPO, GRPO, and GSPO compared across both paradigms.  \n(Gururangan et al., 2020), robustness to dialectal variation within a single language remains comparatively underexplored. Existing work is concentrated heavily on AAVE and code-switching (Blodgett et al., 2016 ; Sap et al., 2019 ; Hofmann et al., 2024), while regional varieties such as Australian, Indian, and Northern British English receive substantially less attention at the alignment stage. Methods that do address dialectal variation—TADA (Held et al., 2023), HyperLoRA (Xiao et al., 2023)—do so via dialect adapters or low-rank components targeting NLU robustness under dialectal input; LoRDD (Srirag et al., 2025a) similarly employs a low-rank dialect adapter for decoder models in a task-specific setting. None of them investigates full-pipeline adaptation across pretraining, fine-tuning, and alignment, nor whether models can produce dialectally appropriate output.  \nDialectal generation (producing a language variety on output, not merely tolerating it on input) is  \na critical part of the problem that current methods do not address. Consider three users chasing a late parcel:  \nIndian English: My parcel is not yet coming; kindly do the needful and prepone the delivery.  \nNorthern British English: Me parcel’snot arrived; summat were","cbCaihpgyXHItX1i","https://ap.wps.com/l/cbCaihpgyXHItX1i","pdf",401670,3,1,18,"English","en",105,"# Abstract\n# Introduction\n## Dialectal robustness vs dialectal generation\n## DiaLLM framework and training pipeline\n## Contributions and findings","[{\"question\":\"What problem does DiaLLM address in dialect adaptation?\",\"answer\":\"DiaLLM targets the gap between dialectal robustness (understanding dialect input) and dialectal generation (producing dialect-appropriate output), which many models fail to do.\"},{\"question\":\"How does DiaLLM adapt models after continual pretraining?\",\"answer\":\"DiaLLM applies two posttraining paradigms—implicit adaptation and explicit adaptation—then evaluates alignment methods such as DPO, GRPO, and GSPO across multiple model families.\"},{\"question\":\"What do the results show about alignment and dialect quality?\",\"answer\":\"Dialectal robustness is mainly driven by continual pretraining and SFT, while alignment changes generation in ways benchmarks miss. Explicit variety-targeted adaptation is reliably preferred, but the strongest dialect-reward-optimizing method is not preferred by human evaluators.\"}]",1784186471,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"diallm-investigating-the-robustness-generation-gap-in-english-dialect-adaptation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/diallm-investigating-the-robustness-generation-gap-in-english-dialect-adaptation/83281/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DiaLLM address in dialect adaptation?","Question",{"text":75,"@type":76},"DiaLLM targets the gap between dialectal robustness (understanding dialect input) and dialectal generation (producing dialect-appropriate output), which many models fail to do.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DiaLLM adapt models after continual pretraining?",{"text":80,"@type":76},"DiaLLM applies two posttraining paradigms—implicit adaptation and explicit adaptation—then evaluates alignment methods such as DPO, GRPO, and GSPO across multiple model families.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about alignment and dialect quality?",{"text":84,"@type":76},"Dialectal robustness is mainly driven by continual pretraining and SFT, while alignment changes generation in ways benchmarks miss. Explicit variety-targeted adaptation is reliably preferred, but the strongest dialect-reward-optimizing method is not preferred by human evaluators.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]