[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86003-en":3,"doc-seo-86003-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86003,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","LayerNorm as Implicit Gain Control in Looped Transformers","LayerNorm within pre-LayerNorm looped transformers functions as an implicit gain controller that links the recurrent block’s local Lipschitz constant inversely to activation scale, making the recurrence Jacobian non-normal. Convergence and stability are shown to depend on the spectral margin rather than an operator-norm bound, with the margin shrinking as carry ρ approaches 1 and some initializations failing to converge to fixed points. Experiments across six tasks, including ablations, indicate linear carry does not drive depth-memory except on axis-aligned per-channel tasks. Results are analytical and validated on a from-scratch CPU-scale implementation.","arXiv :2607 . 1068 1v 1 [ cs .LG] 12 Jul 2026  \nLayerNorm as Implicit Gain Control in Looped Transformers  \nMatthias M. M. Buehlmaier  \nFaculty of Business and Economics, The University of Hong Kong Pokfulam Road, Hong Kong  \nbuehl@hku. hk  \nJuly 12, 2026  \nAbstract  \nIn pre-LayerNorm looped transformers, LayerNorm inside the recurrent block acts as an implicit gain controller: by coupling the block’s local Lipschitz constant inversely to the activation scale, it renders the recurrence Jacobian non-normal—asymptotically contractive at every verified fixed point even where its operator norm exceeds 1—so the true stability budget is the spectral margin, not an operator-norm bound. That margin depletes as the carry ρ → 1, and a minority of initializations never converge to a fixed point at all, so the diagonal carry constraint ρ (¯A) \u003C 1 is necessary but not sufficient for convergence of the full recurrence. Training experiments across six tasks, including a controlled ablation, reveal that the linear carry is not the depth-memory mechanism: gradient descent routes memory through the block’s more expressive nonlinear recurrence and leaves the stability-constrained carry at rest—the carry’s role is stabilization, not memory. We characterize the boundary of this claim: on tasks with axis-aligned per-channel structure, gradient descent does recruit the carry. All results are derived analytically and verified in a from-scratch, CPU-scale implementation; verification at larger scale is needed.  \n1 Introduction  \nLooped transformers—architectures that reuse a single block of transformer weights across multiple depth iterations—decouple reasoning depth from parameter count, offering substantial inferencetime savings. A model with k unique layers iterated T times achieves the effective depth of akT-layer network while storing only the k shared layers’weights. At fixed quality this translates into large parameter savings: Prairie et al. [2026] report matching a Transformer’s quality with roughly half the parameters. This makes the architecture attractive for deployment at scale, where inference costs dominate and parameter efficiency translates directly to serving economics.  \nHowever, the training cost remains comparable to conventional transformers of equivalent effective depth. Each iteration of the shared block contributes to the forward pass and must be backpropagated through, so the FLOP budget scales with effective depth regardless of parameter sharing. The savings are in memory (fewer unique weights) and inference (smaller model to serve), not in the training compute required to learn those weights.  \nThis asymmetry has a practical consequence: looped transformers will be deployed primarily by organizations that can afford frontier-scale training runs—and stability failures during such runs are disproportionately costly. A single training divergence at 70B parameters wastes weeks of compute time and hundreds of thousands of dollars. Understanding the stability properties of the recurrentdepth architecture at small scale, before committing to expensive training runs, is therefore not merely pedagogical but economically necessary.  \nThis paper contributes three previously uncharacterized stability properties of looped transformers with pre-LayerNorm blocks:  \nFirst, LayerNorm inside the recurrent block acts as an implicit gain controller, coupling the block’s local Lipschitz constant inversely to the activation scale: the increment’s Lipschitz constant tracks 1 − ρ (the affine dependence ρ + CF (1 − ρ) is derived, with the constant of proportionality C/F ≈ 2.5 measured near-constant across carries) . This coupling does not pin an operator-norm bound near 1—that bound descends from ≈2 to 1 as ρ → 1, and the measured operator norm itself exceeds 1 at low carry—but it renders the recurrence Jacobian non-normal: across 10 random-weight seeds, every instance that converges to a fixed point has operator norm above 1 there (measured at ρ =","cbCairc9FSgZQsop","https://ap.wps.com/l/cbCairc9FSgZQsop","pdf",648649,1,23,"English","en",105,"# Abstract\n# Introduction\n## Practical motivation for looped transformers\n## Stability properties of pre-LayerNorm looped transformers\n## Experimental evaluation and implementation notes","[{\"question\":\"How does LayerNorm act as an implicit gain controller in pre-LayerNorm looped transformers?\",\"answer\":\"LayerNorm couples the recurrent block’s local Lipschitz constant inversely to the activation scale, inducing non-normal Jacobians and controlling recurrence behavior through this coupling.\"},{\"question\":\"What determines stability in these looped transformers: operator-norm bounds or spectral margin?\",\"answer\":\"Stability budget is governed by the spectral margin (1 − ρspec) rather than operator-norm bounds, and the spectral margin decreases as carry ρ approaches 1.\"},{\"question\":\"Does linear carry provide the depth-memory mechanism during training?\",\"answer\":\"Across cross-channel tasks, gradient descent routes memory through the nonlinear recurrent dynamics and leaves the carry at rest, indicating the linear carry stabilizes rather than stores memory; on axis-aligned per-channel tasks, the claim’s boundary includes scenarios where the carry is recruited.\"}]",1784207697,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"layernorm-as-implicit-gain-control-in-looped-transformers","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/layernorm-as-implicit-gain-control-in-looped-transformers/86003/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does LayerNorm act as an implicit gain controller in pre-LayerNorm looped transformers?","Question",{"text":75,"@type":76},"LayerNorm couples the recurrent block’s local Lipschitz constant inversely to the activation scale, inducing non-normal Jacobians and controlling recurrence behavior through this coupling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What determines stability in these looped transformers: operator-norm bounds or spectral margin?",{"text":80,"@type":76},"Stability budget is governed by the spectral margin (1 − ρspec) rather than operator-norm bounds, and the spectral margin decreases as carry ρ approaches 1.",{"name":82,"@type":73,"acceptedAnswer":83},"Does linear carry provide the depth-memory mechanism during training?",{"text":84,"@type":76},"Across cross-channel tasks, gradient descent routes memory through the nonlinear recurrent dynamics and leaves the carry at rest, indicating the linear carry stabilizes rather than stores memory; on axis-aligned per-channel tasks, the claim’s boundary includes scenarios where the carry is recruited.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]