[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83126-en":3,"doc-seo-83126-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83126,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Pre-Training on Software Engineering Texts: Effects on Domain Adaptation and General-Language Understanding","Generalist and code-focused language models are increasingly used in software engineering, yet evidence on whether they are optimized for SE textual artifacts such as issues and developer discussions is limited. This study compares continual pre-training (CPT) and pre-training from scratch (PTS) on a new SE corpus, evaluating domain adaptation with SELU and general-language understanding with SuperGLUE. With constant-token and compute-matched budgets, CPT yields small, mostly inconclusive domain gains and leaves general understanding unchanged, while PTS incurs large penalties. Practical adaptation guidance and a replication kit are released.","Pre-Training on Software Engineering Texts: Effects on Domain Adaptation and General-Language Understanding  \n1st Fabian C. Pe˜na   \nFaculty of Computer Science and Mathematics University of Passau Passau, Germany [fabiancamilo.penalozano@uni-passau.de](fabiancamilo.penalozano@uni-passau.de)  \n2nd Steffen Herbold  Faculty of Computer Science nd Mathematics University of Passau Passau, Germany [steffen.herbold@uni-passau.de](steffen.herbold@uni-passau.de)  \narXiv :2607 .066 13v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Generalist and code-focused Language Models (LMs) are increasingly applied to software engineering (SE), yet whether they are optimized for understanding SE textual artifacts (e.g., issues, commit messages, developer discussions) remains unclear, as most evidence comes from code-focused benchmarks. We study how to adapt encoder and decoder LMs to SE text, comparing continual pre-training (CPT) against pre-training from scratch (PTS) on a new SE corpus, and evaluating both domain adaptation (SELU) and general-language understanding (SuperGLUE). To keep the comparisons fair, we control pretraining under constant-token and compute-matched budgets. We find that across families and sizes, reusing an existing LM dominates training a domain-native one from scratch: CPT yields small and mostly inconclusive domain gains while leaving general-language understanding essentially unchanged, whereas PTS pays a large and usually decisive penalty on both axes and becomes competitive only for small LMs under a tokenrich budget. We distill these results into practical guidance for adapting LMs to SE text and release our corpus and pre-trained LMs in our replication kit [1].  \nIndex Terms—Software engineering, language models, pretraining, domain adaptation, general-language understanding  \nI. INTRODUCTION  \nIn recent years, Language Models (LM) have been adopted in a growing array of professional tasks, with particularly visible use in software engineering (SE) and other knowledgework activities [2] . Although empirical studies report substantial productivity gains in some controlled settings, these effects are highly heterogeneous and depend on task structure and worker expertise [3], [4] . Frontier-LM development is evolving at a rapid pace, with successive releases introducing improvements in capabilities, efficiency, and alignment [5]–[7]; however, evaluating these improvements usually depends on narrow benchmarks which lack real-world representativeness. Meanwhile, open questions remain regarding reliability, robustness [8], [9], and the extent to which apparent general capabilities transfer to specialized domains. In SE, for example, popular code-focused benchmarks such as SWEbench [10] and Terminal-bench [11] evaluate LMs on issue resolution and CLI handling, respectively, leaving other aspects of the SE practice in the dark.  \nOnly a recent benchmark, SELU [12], enables evaluating language understanding capabilities through the broader SE domain with a cornerstone in SE textual artifacts specific to code-adjacent metadata, pure non-code settings, and developer communications. By running SELU on a range of LMs, the authors found that those that might be considered outdated (e.g., GPT-2 [13]) may outperform larger, more recent and even code-focused LMs (e.g., Llama 3.2 [14] and StarCoder2 [15]) . We hypothesize that this improvement stems from optimizing LMs for SE textual artifacts from a foundational stage (i.e., pre-training), feeding them high-quality data aligned with this purpose.  \nBased on this hypothesis, we systematically study domain adaptation by applying continual pre-training (CPT) [16]–[18] to a range of open-source LMs on a new corpus of SE text collected from sources such as GitHub, Stack Overflow, Jira, andarXiv. Similarly to CPT, we also pre-train LMs from scratch (PTS), i.e., not starting from an existing LM but from a new one with randomly initialized weights. Throughout, domain adaptation refers to adapting LM","cbCaibSImLfoj9LQ","https://ap.wps.com/l/cbCaibSImLfoj9LQ","pdf",292701,1,12,"English","en",105,"# Abstract\n# I. Introduction\n# II. Background and Related Work","[{\"question\":\"What question does the study address about software engineering language models?\",\"answer\":\"Whether language models are optimized for understanding software engineering textual artifacts such as issues, commit messages, and developer discussions, beyond evidence from code-focused benchmarks.\"},{\"question\":\"How are continual pre-training (CPT) and pre-training from scratch (PTS) compared in the paper?\",\"answer\":\"The paper adapts encoder and decoder language models to SE text using CPT versus PTS on a new SE corpus, while controlling token and compute budgets to keep comparisons fair.\"},{\"question\":\"What do the results say about CPT versus PTS for domain adaptation and general-language understanding?\",\"answer\":\"CPT produces small and mostly inconclusive domain gains with essentially unchanged general-language understanding, whereas PTS shows large and usually decisive penalties on both metrics and only becomes competitive for small models under a token-rich budget.\"}]",1784185459,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"pre-training-on-software-engineering-texts-effects-on-domain-adaptation-and-general-language-understanding","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/pre-training-on-software-engineering-texts-effects-on-domain-adaptation-and-general-language-understanding/83126/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the study address about software engineering language models?","Question",{"text":75,"@type":76},"Whether language models are optimized for understanding software engineering textual artifacts such as issues, commit messages, and developer discussions, beyond evidence from code-focused benchmarks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are continual pre-training (CPT) and pre-training from scratch (PTS) compared in the paper?",{"text":80,"@type":76},"The paper adapts encoder and decoder language models to SE text using CPT versus PTS on a new SE corpus, while controlling token and compute budgets to keep comparisons fair.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results say about CPT versus PTS for domain adaptation and general-language understanding?",{"text":84,"@type":76},"CPT produces small and mostly inconclusive domain gains with essentially unchanged general-language understanding, whereas PTS shows large and usually decisive penalties on both metrics and only becomes competitive for small models under a token-rich budget.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]