[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82064-en":3,"doc-seo-82064-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82064,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","How are linear representations learned Exact solutions to the dynamics of abstraction","Linear directions in representation space often encode concepts in artificial and biological neural networks, motivating the linear representation hypothesis used by many interpretability and control methods. This work develops a dynamical framework for how concept directions align during training, termed “abstraction.” In a minimal linear network, exact solutions give the full abstraction trajectory and show how data/target geometry, depth, and initialization scale shape end-of-learning abstraction. For nonlinear networks, erf matches the linear theory, while ReLU depends more on input geometry and exhibits an attenuation law, validated in public models and applied to improve linear probe generalization in LLMs.","arXiv :2607 .08843v 1 [ cs .LG] 9 Jul 2026  \nHow are linear representations learned? Exact solutions to the dynamics of abstraction  \nWilliam W. Yang 1 Andrew M. Saxe∗ , 1 ,2 Peter E. Latham∗ , 1  \nAbstract  \nIn artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist after training, the dynamics of how they emerge during training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training  \n– a process we call “abstraction”. In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning,(ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.  \n1 Introduction  \nA recurring empirical observation in deep neural networks is that high-level concepts often behave like approximately linear directions in representation space [1, 2] . This idea has recently been formalized as the linear representation hypothesis (LRH) [3, 4, 5] . To use a classical example, alinear representation of the “gender” concept would imply that the concept vectors vking − vqueen and vman − vwoman are approximately parallel. A closely related idea exists in neuroscience, where a concept is said to be represented in an “abstract” format when concept vectors are highly aligned across contexts [6, 7, 8, 9, 10, 11, 12, 13] . Borrowing that neuroscience terminology, we use the term “abstraction” to refer to the alignment of concept vectors (e.g. the cosine similarity between vking − vqueen and vman − vwoman ) during training.  \nUnderstanding how abstract representations arise is important in both deep learning and neuroscience. Many AI interpretability and control methods assume existence of abstract representations by using linear probes to extract concept directions for interpretability and detection [14, 15] or model steering [16, 17] . In neuroscience, abstraction is thought to support the brain’s ability to adapt to changing environments via out-of-distribution and compositional generalization [6, 10, 18] .  \n*Co-senior authors; equal contribution.  \n1 Gatsby Computational Neuroscience Unit, University College London, London W1T 4JG, United Kingdom.  \n2 Sainsbury Wellcome Centre, University College London, London W1T 4JG, United Kingdom.  \nPreprint.  \nHowever, our understanding of how abstraction arises remains limited. Previous theoretical work has argued that abstraction becomes perfect (i.e. cosine similarity = 1) once trained to convergence [4], or at the global minima of the loss landscape [19] . This leaves open a simple question: what are the dynamics of abstraction during training? Empirically, in real-world settings abstraction often exhibits nontrivial dynamics, can be non-monotonic, and seldom reach perfect abstraction as predicted in simple theories; for example, see ","cbCaik8dtbGfd58j","https://ap.wps.com/l/cbCaik8dtbGfd58j","pdf",3654828,1,81,"English","en",105,"# Introduction\n## Linear representation hypothesis and abstraction\n## Motivation and central questions\n## Main contributions","[{\"question\":\"What does the paper mean by “abstraction” during training?\",\"answer\":\"Abstraction refers to the alignment of concept vectors over training, measured for example by cosine similarity between concept-difference directions across contexts.\"},{\"question\":\"What does the paper’s minimal linear network analysis provide?\",\"answer\":\"It derives exact solutions for the full trajectory of abstraction during training, identifying how data/target geometry, network depth, and initialization scale determine abstraction outcomes.\"},{\"question\":\"How do erf and ReLU networks differ in abstraction dynamics?\",\"answer\":\"The theory indicates erf networks approximate the linear case, while ReLU abstraction depends less on target geometry and more on input geometry. The paper also proves that nonlinearities attenuate abstraction in activations relative to preactivations.\"}]",1784177968,204,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"how-are-linear-representations-learned-exact-solutions-to-the-dynamics-of-abstraction","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/how-are-linear-representations-learned-exact-solutions-to-the-dynamics-of-abstraction/82064/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper mean by “abstraction” during training?","Question",{"text":75,"@type":76},"Abstraction refers to the alignment of concept vectors over training, measured for example by cosine similarity between concept-difference directions across contexts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper’s minimal linear network analysis provide?",{"text":80,"@type":76},"It derives exact solutions for the full trajectory of abstraction during training, identifying how data/target geometry, network depth, and initialization scale determine abstraction outcomes.",{"name":82,"@type":73,"acceptedAnswer":83},"How do erf and ReLU networks differ in abstraction dynamics?",{"text":84,"@type":76},"The theory indicates erf networks approximate the linear case, while ReLU abstraction depends less on target geometry and more on input geometry. The paper also proves that nonlinearities attenuate abstraction in activations relative to preactivations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]