[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120966-en":3,"doc-seo-120966-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120966,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Unnatural Algorithms in Machine Learning - Paper","Natural gradient descent remains invariant under network reparameterizations in the small learning-rate limit, yielding robust optimization behavior. The work studies a broader class of discrete optimization schemes that approximate natural transformations between optimizer-state functors over the diffeomorphism group. These algorithms exhibit improved efficiency for poorly parameterized networks, with generated parameter flows becoming invariant under smooth reparameterizations via equivariant maps. The paper generalizes naturality beyond group equivariance (including non-invertible projections), enabling comparisons across non-isomorphic architectures and analysis via inverse limits. It presents a method to enforce this naturality and finds that many popular training algorithms are unnatural.","arXiv :2312 .04739v1 [ stat .ML] 7 Dec 2023  \nUnnatural Algorithms in Machine Learning  \nChristian Goodbrake [christian.goodbrake@oden.utexas.edu](christian.goodbrake@oden.utexas.edu)  \nOden Institute for Computational Engineering and Sciences University of Texas  \nAustin, TX 78712, USA  \nEditor:  \nAbstract  \nNatural gradient descent has a remarkable property that in the small learning rate limit, it displays an invariance with respect to network reparameterizations, leading to robust training behavior even for highly covariant network parameterizations. We show that optimization algorithms with this property can be viewed as discrete approximations of natural transformations from the functor determining an optimizer’s state space from the di􀀋eomorphism group if its con􀀌guration manifold, to the functor determining that state space’s tangent bundle from this group. Algorithms with this property enjoy greater e􀀎ciency when used to train poorly parameterized networks, as the network evolution they generate is approximately invariant to network reparameterizations. More speci􀀌cally, the 􀀍ow generated by these algorithms in the limit as the learning rate vanishes is invariant under smooth reparameterizations, the respective 􀀍ows of the parameters being determined byequivariant maps. By casting this property a natural transformation, we allow for generalizations beyond equivariance with respect to group actions; this framework can account for non-invertible maps such as projections, creating a framework for the direct comparison of training behavior across non-isomorphic network architectures, and the formal examination of limiting behavior as network size increases by considering inverse limits of these projections, should they exist. We introduce a simple method of introducing this naturality more generally and examine a number of popular machine learning training algorithms, 􀀌nding that most are unnatural.  \nKeywords: natural gradient descent, natural transformation, reparameterization invariance, group equivariance  \n1 Introduction  \nDrawing from the world of information geometry, the “natural gradient descent” (Amari, 1998, 2010) has been proposed and widely studied as an alternative to the traditional gradient descent algorithm for training neural networks. In this approach, the Fisher information matrix is used to endow the network’s parameter space with a Riemannian structure. With this added structure, the step taken at each update is in the direction of the steepest descent relative to the space of realizable distributions, rather than with respect to the network’s parameters. The Fisher information matrix is used to “correct” the ordinary gradient to account for the the geometry of the network’s parameters, and it has been shown that this can improve convergence by accounting for covariance between network parameters analogous to signal whitening (Sohl-Dickstein, 2012) . Remarkably, this structure is invariant under reparameterizations of the network, as the components of the Fisher information matrix  \n©2023 Christian Goodbrake.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/. Attribution)[. Attribution](https://creativecommons.org/licenses/by/4.0/. Attribution) requirements are provided  \nat [http://jmlr.org/papers/v/.html](http://jmlr.org/papers/v/.html).  \nGoodbrake  \ntransform in a manner complementary to the components of the gradient under reparameterizations, leading to robust training behavior for natural gradient descent even in poorly parameterized networks (Pascanu and Bengio, 2013) . This structure has been studied in much depth, leading to the family of algorithms derived from the information-geometric optimization approach (Ollivier et al., 2017) .  \nVariants of this approach have been developed along numerous lines. Martens (2020) notes that the natural gradient descent can be viewed as a second order method, placing it alongside other ","cbCaitWDc2cOXLhu","https://ap.wps.com/l/cbCaitWDc2cOXLhu","pdf",231439,1,20,"English","en",105,"# Introduction\n## Natural gradient descent from information geometry\n## Variants and connections to second-order optimization\n## Goal: generalizing invariance via naturality","[{\"question\":\"What key property motivates the paper’s study of optimization algorithms?\",\"answer\":\"Natural gradient descent becomes invariant under network reparameterizations in the small learning-rate limit, producing robust training behavior.\"},{\"question\":\"How does the paper connect optimization to natural transformations?\",\"answer\":\"It models algorithms with the invariance property as discrete approximations of natural transformations between functors that define the optimizer’s state space and its tangent bundle over a diffeomorphism group.\"},{\"question\":\"What does the framework enable beyond standard group equivariance?\",\"answer\":\"It can handle non-invertible maps such as projections, supporting direct comparison of training behavior across non-isomorphic network architectures and examination of limiting behavior as network size grows via inverse limits.\"}]","Unnatural Algorithms in Machine Learning - Paper | PDF",1785733096,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"unnatural-algorithms-in-machine-learning-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/unnatural-algorithms-in-machine-learning-paper/120966/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key property motivates the paper’s study of optimization algorithms?","Question",{"text":75,"@type":76},"Natural gradient descent becomes invariant under network reparameterizations in the small learning-rate limit, producing robust training behavior.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper connect optimization to natural transformations?",{"text":80,"@type":76},"It models algorithms with the invariance property as discrete approximations of natural transformations between functors that define the optimizer’s state space and its tangent bundle over a diffeomorphism group.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the framework enable beyond standard group equivariance?",{"text":84,"@type":76},"It can handle non-invertible maps such as projections, supporting direct comparison of training behavior across non-isomorphic network architectures and examination of limiting behavior as network size grows via inverse limits.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]