[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84505-en":3,"doc-seo-84505-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84505,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics","Studies feed-forward ReLU networks with fixed readout and quadratic loss by rewriting gradient descent as collective activation-field dynamics and conjugate (backpropagated) fields on the training set. Working to first order in the learning rate within a fixed ReLU activation chamber, the note derives explicit one-, two-, and three-hidden-layer cases and then an arbitrary-depth recursion. It identifies how weight-induced pullback Gram metrics first enter the conjugate-field dynamics, and how cut-wise push-forward/pullback transport reconstructs layerwise residual kernels.","arXiv :2606 .09744v4 [ cs .LG] 11 Jul 2026  \nLearning Dynamics Reveal a Hierarchy of Weight-Induced  \nLayerwise Gram Metrics  \nClaudio Nordio  \nDraft research note∗  \nJuly 14, 2026  \nAbstract  \nWe study feed-forward ReLU networks with fixed readout and quadratic loss, and rewrite gradient descent as a collective dynamics of activation fields and conjugate fields on the training set. Working to first order in the learning rate inside a fixed activation chamber, we derive explicitly the one-, two-and three-hidden-layer cases, and then give the arbitrary-depth recursion. For one hidden layer the activation dynamics closes directly and the residual update is governed by the product of an input Gram matrix and a co-activation/backpropagation Gram matrix. For two hidden layers a conjugate field is required, but no nontrivial pullback Gram metric has yet appeared. For three hidden layers the first weight-induced pullback Gram metric enters the conjugate-field dynamics. At arbitrary depth, activation variations propagate forward through a recursive response operator Uℓαβ , while conjugate-field variations propagate backward through an effective transport operator Mαβℓ . Their contractions reconstruct alayerwise residual kernel  \nL  \nKα(Lβ) = X Qα(ℓ1) Sα(ℓβ) .  \nℓ=1  \nThe resulting description exposes a duality between push-forward and pullback transport across every layer cut, and identifies the first Gram metrics as the lowest nontrivial terms ina broader hierarchy of activation-conditioned transport operators. We deliberately stop atthe level of collective fields, conjugate fields, residual kernels and cut-wise transport metrics, leaving the later tensorial geometric formulation outside the scope of this paper.  \n1 Introduction  \nGradient descent in neural networks is usually formulated as a dynamical system in weight space. The trainable variables are the weight matrices, while activations, predictions and residuals are derived quantities. This is the natural optimization viewpoint, but it is not always the most transparent way to describe the internal organization produced by learning.  \nThis paper develops a complementary formulation for feed-forward ReLU networks with fixed readout. The central idea is to rewrite gradient descent as a dynamics of collective fields defined on the training set: the activation fields uαℓ, their conjugate or backpropagated fields bαℓ, and the induced transport operators through which variations of these fields propagate across depth. The formulation is local in training time: throughout the derivation we work to first order in the learning rate η, inside a fixed ReLU activation chamber. The masks are therefore held fixed during an infinitesimal update.  \nThe motivation comes from the simplest case. With one hidden layer, the weight update can be eliminated from the activation dynamics. The residuals obey a closed kernel equation  \n∗ Circulated for discussion and feedback; comments are welcome at [c.nordio@gmail.com](c.nordio@gmail.com).  \nwhose kernel factorizes into an input overlap and a co-activation/backpropagation overlap. This resembles the spirit of NTK descriptions [2], but the emphasis here is different: rather than freezing the kernel in a limiting regime, we track how the finite-depth collective fields and their transport operators appear directly from gradient descent.  \nIncreasing depth reveals a structured hierarchy. For two hidden layers one must introduce conjugate fields, since updating the second weight matrix changes the backpropagated field atthe first hidden layer. For three hidden layers, a new object appears: a weight-induced pullback Gram metric of the form  \nGαβℓ = 􀀐 W (ℓ+1)􀀑 T D1 W (ℓ+1) , (1)  \nwhere D1 = Aαℓ+1 Aβℓ+1 is a co-activation projector. Deeper networks generate further iterates of the same pullback transport mechanism. Thus the first Gram metric is not an isolated curiosity: it is the first nontrivial member of a hierarchy.  \nThe goal of this paper is to make thi","cbCaiidtcfFJ18XJ","https://ap.wps.com/l/cbCaiidtcfFJ18XJ","pdf",371586,1,11,"English","en",105,"# Abstract\n# Introduction\n# Setup and Fixed-Chamber ReLU Dynamics\n## Definitions of activations, masks, and forward map\n## Fixed readout, predictions, residuals, and loss\n## Conjugate fields via backward transport","[{\"question\":\"What is the core reformulation of gradient descent proposed here?\",\"answer\":\"Gradient descent is rewritten as collective dynamics of activation fields and their conjugate (backpropagated) fields defined on the training set, rather than as a dynamical system in weight space.\"},{\"question\":\"What does working “to first order in the learning rate” mean in this note?\",\"answer\":\"The derivation keeps the learning-rate-dependent update only at first order and assumes a fixed ReLU activation chamber, so the ReLU masks remain unchanged during an infinitesimal update.\"},{\"question\":\"When do weight-induced pullback Gram metrics first appear?\",\"answer\":\"They do not appear nontrivially at the two-hidden-layer level, but the first weight-induced pullback Gram metric enters the conjugate-field dynamics for three hidden layers, starting a broader hierarchy at deeper depths.\"}]",1784196179,28,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"learning-dynamics-reveal-a-hierarchy-of-weight-induced-layerwise-gram-metrics","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/learning-dynamics-reveal-a-hierarchy-of-weight-induced-layerwise-gram-metrics/84505/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core reformulation of gradient descent proposed here?","Question",{"text":75,"@type":76},"Gradient descent is rewritten as collective dynamics of activation fields and their conjugate (backpropagated) fields defined on the training set, rather than as a dynamical system in weight space.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does working “to first order in the learning rate” mean in this note?",{"text":80,"@type":76},"The derivation keeps the learning-rate-dependent update only at first order and assumes a fixed ReLU activation chamber, so the ReLU masks remain unchanged during an infinitesimal update.",{"name":82,"@type":73,"acceptedAnswer":83},"When do weight-induced pullback Gram metrics first appear?",{"text":84,"@type":76},"They do not appear nontrivially at the two-hidden-layer level, but the first weight-induced pullback Gram metric enters the conjugate-field dynamics for three hidden layers, starting a broader hierarchy at deeper depths.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]