[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86532-en":3,"doc-seo-86532-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86532,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Domain-Aware Scaling Laws Uncover Data Synergy","Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings show that combining datasets from different domains yields nontrivial interactions. Code can improve mathematical reasoning, while certain mixtures introduce interference that reduces performance. This work formalizes and quantifies data synergy in language model pretraining, estimating direct domain-to-benchmark synergy and second-order domain-domain synergy, improving predictive accuracy and enabling stable ranking of mixture performance.","arXiv :2607 . 1 1052v 1 [ cs .LG] 13 Jul 2026  \nDomain-Aware Scaling Laws Uncover Data Synergy  \nKimia Hamidieh1 ∗ Lester Mackey2 David Alvarez-Melis2,3  \n1MIT CSAIL, 2 Microsoft Research, 3 Harvard University  \nAbstract  \nMachine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential.  \nEmpirical findings repeatedly show that combining datasets from different domains yields nontrivial interactions. For instance, adding code improves mathematical reasoning, while certain mixtures introduce interference that reduces model performance. We refer to these effects collectively as data synergy, where the contribution of multiple domains exceeds or falls short of the sum of their isolated contributions. In this work, we formalize and quantify data synergy in language model pretraining. Leveraging observational variation across open-weight LLMs with diverse pretraining mixtures, we estimate both direct domain-to-benchmark synergy (how one domain contributes to performance on another) and a second-order domain-domain synergy (capabilities that require co-occurrence of multiple domains) . Our framework improves predictive accuracy over domain-agnostic scaling laws and recovers stable synergy estimates. We validate these estimates by training models on predicted optimal and predicted anti-optimal mixtures and confirm that our synergy estimates correctly predict performance rankings.  \n1 Introduction  \nRecent improvements in Large Language Models (LLMs) are strongly shaped by their pretraining data distribution (Li et al., 2024a; Soldaini et al., 2024), yet prior work abstracts away composition and reduces data to an undifferentiated token count (Kaplan et al., 2020; Hoffmann et al., 2022) . In practice, however, pretraining corpora are mixtures of data from different domains (e.g., web, books, code, math) and small shifts in composition can significantly impact model capabilities (Ye et al., 2024; Liu et al., 2024) . In particular, empirical studies repeatedly report interactions between data domains that are systematic rather than anecdotal: code pretraining can improve mathematical (Lu et al., 2024; Azerbayev et al., 2023) and logical (Yu et al., 2024) reasoning, math and code mixtures often outperform either alone (Aryabumi et al., 2024), while other combinations lead to interference and degrade performance on certain tasks (Zheng et al., 2024; Li et al., 2024b; Gu et al., 2024) . These observations suggest that tokens are not interchangeable. What matters is not only how much data we train on, but also what kinds of data are combined.  \nWe refer to these interactions collectively as data synergy. We distinguish two complementary forms (Figure 1). The first is domain→benchmark synergy: the extent to which pretraining data from one domain helps or hurts performance on a given benchmark, beyond what would be expected from total token count. The second is domain-domain synergy: nonadditive effects that arise when two domains co-occur in the training mixture, so that their joint contribution is greater or smaller than the sum of their isolated effects. The former captures the effect of a training domain on a given benchmark or evaluation domain, and the latter captures higher-order complementarities or interference internal to the pretraining corpus itself.  \n∗ [Correspondence to](Correspondence to hamidieh@mit.edu)[ hamidieh@mit.edu](Correspondence to hamidieh@mit.edu).  \nPretraining Domain → Benchmark Synergy  \nCoding Benchmark  \nPrompt: def truncate_number(number:  float) -> float:  \n\"\"\" Given a positive floating point   \nnumber …>>> truncate_number(3.5) 0.5 “\"\"  \nResponse : Here is the code for …  \nγmath > 0 γweb \u003C 0  \nSecond-Order Domain-Domain Synergy  \nSynergy σweb,QA  \n“Bonus” Tokens  \nData Size D  \nEffective Data Size D eff  \nFigure 1: Two forms of data synergy: Pretraining corpora are mixtures of different data domains (Web, Books, Math, Q&A, Scien","cbCaisJfLT8W8i9M","https://ap.wps.com/l/cbCaisJfLT8W8i9M","pdf",1663319,2,1,25,"English","en",105,"# Abstract\n# Introduction\n## Data synergy in pretraining corpora","[{\"question\":\"What is “data synergy” in language model pretraining?\",\"answer\":\"Data synergy refers to situations where multiple data domains contribute more (or less) to performance than the sum of their isolated contributions. It captures nonadditive interactions across domains during pretraining.\"},{\"question\":\"How does the paper quantify data synergy?\",\"answer\":\"It formalizes domain-aware scaling laws and estimates two components: domain→benchmark synergy and second-order domain-domain synergy. Estimates are derived from observational variation across open-weight LLMs with diverse pretraining mixtures.\"},{\"question\":\"How is the proposed synergy framework validated?\",\"answer\":\"Models are trained on predicted optimal and predicted anti-optimal mixtures. The resulting performance rankings are used to confirm that synergy estimates correctly predict which mixtures work best.\"}]",1784212448,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"domain-aware-scaling-laws-uncover-data-synergy","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/domain-aware-scaling-laws-uncover-data-synergy/86532/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is “data synergy” in language model pretraining?","Question",{"text":75,"@type":76},"Data synergy refers to situations where multiple data domains contribute more (or less) to performance than the sum of their isolated contributions. It captures nonadditive interactions across domains during pretraining.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper quantify data synergy?",{"text":80,"@type":76},"It formalizes domain-aware scaling laws and estimates two components: domain→benchmark synergy and second-order domain-domain synergy. Estimates are derived from observational variation across open-weight LLMs with diverse pretraining mixtures.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed synergy framework validated?",{"text":84,"@type":76},"Models are trained on predicted optimal and predicted anti-optimal mixtures. The resulting performance rankings are used to confirm that synergy estimates correctly predict which mixtures work best.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]