[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-203809-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-203809-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","sudden-drops-in-the-loss-syntax-acquisition-phase-transitions-and-simplicity-bias-in-mlms","SUDDEN DROPS IN THE LOSS: SYNTAX ACQUISITION, PHASE TRANSITIONS, AND SIMPLICITY BIAS IN MLMS","","Most interpretability research in NLP studies a fully trained model, but key insights may require examining the training trajectory. This paper presents a case study of syntax acquisition in masked language models, focusing on Syntactic Attention Structure (SAS), an emerging property where Transformer heads align with specific syntactic relations. The work identifies a brief pretraining window when SAS abruptly appears with a steep loss drop, enabling later linguistic capabilities.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/sudden-drops-in-the-loss-syntax-acquisition-phase-transitions-and-simplicity-bias-in-mlms/203809/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/sudden-drops-in-the-loss-syntax-acquisition-phase-transitions-and-simplicity-bias-in-mlms/203809.png","ImageObject",300,407,{"name":42,"@type":43},"Aria Callaghan","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-07","2026-09-04",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",7,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What is Syntactic Attention Structure (SAS) in masked language models?","Question",{"text":63,"@type":64},"SAS is a naturally emerging behavior where specific Transformer attention heads focus on particular syntactic relations, reflecting internal syntactic structure.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"When does the paper observe SAS and what happens to the loss at that time?",{"text":68,"@type":64},"The paper identifies a brief window in pretraining during which SAS abruptly acquires, coinciding with a steep drop in loss.",{"name":70,"@type":61,"acceptedAnswer":71},"How do the authors test whether SAS is necessary for linguistic capabilities?",{"text":72,"@type":64},"They manipulate SAS during training and show that SAS is required for developing grammatical capabilities, while temporarily suppressing SAS improves model quality and convergence.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},203809,1788564027,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,110,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":107,"show_sort_weight":108,"slug":109},"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},8,30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":111,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},962084926284,"https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0","arXiv :2309 .07311v6 [ cs .CL] 21 Mar 2025  \nSUDDEN DROPS IN THE LOSS: SYNTAX ACQUISITION , PHASE TRANSITIONS , AND SIMPLICITY BIAS IN MLMS  \nAngelica Chen 1 Ravid Shwartz-Ziv 1 Kyunghyun Cho 1 ,2 ,3 Matthew L. Leavitt4 Naomi Saphra5  \n{angelica.chen, ravid.shwartz.ziv, [kyunghyun.cho}@nyu.edu](kyunghyun.cho}@nyu.edu)[ ](kyunghyun.cho}@nyu.edu)[matthew@datologyai.com](matthew@datologyai.com) [nsaphra@fas.harvard.edu](nsaphra@fas.harvard.edu)  \n1NYU 2 Genentech 3 CIFAR LMB 4 DatologyAI 5 Kempner Institute, Harvard  \nABSTRACT  \nMost interpretability research in NLP focuses on understanding the behavior and features of a fully trained model. However, certain insights into model behavior may only be accessible by observing the trajectory of the training process. We present a case study of syntax acquisition in masked language models (MLMs) that demonstrates how analyzing the evolution of interpretable artifacts throughout training deepens our understanding of emergent behavior. In particular, we study Syntactic Attention Structure (SAS), a naturally emerging property of MLMs wherein specific Transformer heads tend to focus on specific syntactic relations.  \nWe identify a brief window in pretraining when models abruptly acquire SAS, concurrent with a steep drop in loss. This breakthrough precipitates the subsequent acquisition of linguistic capabilities. We then examine the causal role of SAS by manipulating SAS during training, and demonstrate that SAS is necessary for the development of grammatical capabilities. We further find that SAS competes with other beneficial traits during training, and that briefly suppressing SAS improves model quality. These findings offer an interpretation of a real-world example of both simplicity bias and breakthrough training dynamics.  \n1 INTRODUCTION  \nWhile language model training usually leads to smooth improvements in loss over time (Kaplan et al., 2020), not all knowledge emerges uniformly. Instead, language models acquire different capabilities at different points in training. Some capabilities remain fixed (Press et al., 2023), while others decline (McKenzie et al., 2022), as a function of dataset size or model capacity. Certain capabilities even exhibit abrupt improvements—this paper focuses on such discontinuous dynamics, which are often called breakthroughs (Srivastava et al., 2022), emergence (Wei et al., 2022), breaks (Caballero et al., 2023), or phase transitions (Olsson et al., 2022) . The interpretability literature rarely illuminates how these capabilities emerge, in part because most analyses only examine the final trained model. Instead, we consider developmental analysis as a complementary explanatory lens.  \nTo better understand the role of interpretable artifacts in model development, we analyze and manipulate these artifacts during training. We focus on a case study of Syntactic Attention Structure (SAS), a model behavior thought to relate to grammatical structure. By measuring and controlling the emergence of SAS, we deepen our understanding of the relationship between the internal structural traits and extrinsic capabilities of masked language models (MLMs) .  \nSAS occurs when a model learns specialized attention heads that focus on a word’s syntactic neighbors. This behavior emerges naturally during conventional MLM pre-training (Clark et al., 2019; Voita et al., 2019; Manning et al., 2020) . We observe an abrupt spike in SAS at a consistent point in training, and explore its impact on MLM capabilities by manipulating SAS during training. Our observations paint a picture of how interpretability artifacts may represent simplicity biases that compete with other learning strategies during MLM training. In summary, our main contributions are:  \n• Monitoring latent syntactic structure (defined in Section 2.1) throughout training, we identify (Section 4.1) a precipitous loss drop composed of multiple phase transitions (defined in Section 2.3)  \nrelating to various linguistic abi","cbCaitqgaqvZpuMf","https://ap.wps.com/l/cbCaitqgaqvZpuMf","pdf",1324195,32,"English","# Abstract\n# Introduction\n# Methods\n## Syntactic Attention Structure","[{\"question\":\"What is Syntactic Attention Structure (SAS) in masked language models?\",\"answer\":\"SAS is a naturally emerging behavior where specific Transformer attention heads focus on particular syntactic relations, reflecting internal syntactic structure.\"},{\"question\":\"When does the paper observe SAS and what happens to the loss at that time?\",\"answer\":\"The paper identifies a brief window in pretraining during which SAS abruptly acquires, coinciding with a steep drop in loss.\"},{\"question\":\"How do the authors test whether SAS is necessary for linguistic capabilities?\",\"answer\":\"They manipulate SAS during training and show that SAS is required for developing grammatical capabilities, while temporarily suppressing SAS improves model quality and convergence.\"}]","SUDDEN DROPS IN THE LOSS: SYNTAX ACQUISITION, PHASE TRANSITIONS, AND SIMPLICITY BIAS IN MLMS | PDF",81]