[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125646-en":3,"doc-seo-125646-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125646,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Engineering flexible machine learning systems by traversing functionally-invariant paths - Differential geometry framework","Transformers are the foundation of modern neural networks for language and vision, yet the mathematics that govern adaptation without losing prior knowledge remains unclear. This work proposes a differential geometry framework, functionally-invariant paths (FIP), treating neural weights as a curved Riemannian manifold. Adaptation is defined as geodesic movement in weight space aligned with secondary objectives, using the metric spectrum to identify low-rank subspaces. With modest computation, FIP achieves competitive results on continual learning and sparsification across BERT, vision transformers, and CNNs.","arXiv :2205 .00334v4 [ cs .LG] 3 Sep 2023  \nEngineering flexible machine learning systems by traversing functionally-invariant paths  \nGuruprasad Raghavan 1 , Bahey Tharwat2 ,Surya Narayanan Hari 1 , Dhruvil Satani 1 , Matt Thomson 1∗  \n1 Department of Biology and Biological Engineering, Caltech 1200 E California Blvd, Pasadena, CA 91125  \n2 Alexandria University Alexandria, Egypt  \n∗ To whom correspondence should be addressed; [mthomson@caltech.edu](mthomson@caltech.edu)  \n[graghava@caltech.edu](graghava@caltech.edu)  \nAbstract  \nTransformers have emerged as the state ofthe art neural network architecture for natural language processing and computer vision. In the foundation model paradigm, large transformer models (BERT, GPT3/4, Bloom, ViT) are pre-trained on self-supervised tasks such as word or image masking, and then, adapted through fine-tuning for downstream user applications including instruction following and Question Answering. While many approaches have been developed for model fine-tuning including low-rank weight update strategies (eg. LoRA), underlying mathematical principles that enable network adaptation without knowledge loss remain poorly understood. Here, we introduce a differential geometry framework, functionally invariant paths (FIP), that provides flexible and continuous adaptation of neural networks for a range of machine learning goals and network sparsification objectives. We conceptualize the weight space of a neural network as a curved Riemannian manifold equipped with a metric tensor whose spectrum defines low rank subspaces in weight space that accommodate network adaptation without loss of prior knowledge. We formalize adaptation as movement along a geodesic path in weight space while searching for networks that accommodate secondary objectives. With modest computational resources, the FIP algorithm achieves comparable to state of the art performance on continual learning and sparsification tasks for language models (BERT), vision transformers (ViT, DeIT), and the CNNs. Broadly, we conceptualize a neural network as a mathematical object that can be iteratively transformed into distinct configurations by the path-sampling algorithm to define a sub-manifold of weight space that can be harnessed to achieve user goals.  \nIntroduction  \nTransformers are now the state of the art machine learning paradigm for natural language understanding, computer vision, biological sequence analysis, and context sensitive reasoning tasks [19, 54, 5, 7] . As models have scaled in parameter number, models trained on generic masking tasks have exhibited emergent behaviors including zero-shot task performance, generalized reasoning, and instruction following [7, 11, 43] . In the ‘foundation model’ paradigm [5], transformers with 108 to 1012 parameters are trained over large data sets on self-supervised tasks such as masked language modeling, causal language modeling, or image masking [22, 43, 12, 11] . Following self supervised training, models can be adapted to increase performance on specific applications including ques-  \ntion/answer, instruction following, distillation of financial or medical documents, and sparsified or quantized to reduce memory requirements and inference speeds in deployment environments.  \nDue to the central role of model adaptation for transformer optimization and deployment, many algorithms have emerged for updating model weights to increase performance without experiencing a catastrophic loss of the knowledge gained during self-supervised pre-training. For example, LoRA (Low Rank Adaptation) exhibits impressive performance on fine-tuning of billion parameter models through low rank weight updates enforced by matrix factorization [24] . Empirical experiments with LoRA find that low-rank updates enable fine tuning of models with 100B parameters in benchmark tasks. However, the performance of fine-tuning and sparsification frameworks remains predominantly grounded in empirical results. The machin","cbCaimWSpLiHESJ2","https://ap.wps.com/l/cbCaimWSpLiHESJ2","pdf",4470334,1,22,"English","en",105,"# Abstract\n# Introduction\n## Transformers and foundation models\n## Weight adaptation and fine-tuning\n## Geometric view of weight space\n## Problem setting: preventing knowledge loss","[{\"question\":\"What is functionally-invariant paths (FIP) in this paper?\",\"answer\":\"FIP is a differential geometry framework that enables flexible, continuous neural network adaptation by traversing geodesic paths in weight space while aiming to maintain functional invariance.\"},{\"question\":\"How does the paper model neural network adaptation mathematically?\",\"answer\":\"It treats the weight space as a curved Riemannian manifold with a metric tensor whose spectrum defines low-rank subspaces, and formalizes adaptation as movement along geodesic paths aligned with secondary objectives.\"},{\"question\":\"What tasks does the FIP algorithm target and how does it perform?\",\"answer\":\"FIP targets continual learning and network sparsification objectives and reports performance comparable to state of the art on language models (e.g., BERT), vision transformers, and CNNs with modest computation.\"}]","Engineering flexible machine learning systems by traversing functionally-invariant paths - Differential geometry framework | PDF",1785900399,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"engineering-flexible-machine-learning-systems-by-traversing-functionally-invariant-paths-differential-geometry-framework","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/engineering-flexible-machine-learning-systems-by-traversing-functionally-invariant-paths-differential-geometry-framework/125646/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is functionally-invariant paths (FIP) in this paper?","Question",{"text":75,"@type":76},"FIP is a differential geometry framework that enables flexible, continuous neural network adaptation by traversing geodesic paths in weight space while aiming to maintain functional invariance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper model neural network adaptation mathematically?",{"text":80,"@type":76},"It treats the weight space as a curved Riemannian manifold with a metric tensor whose spectrum defines low-rank subspaces, and formalizes adaptation as movement along geodesic paths aligned with secondary objectives.",{"name":82,"@type":73,"acceptedAnswer":83},"What tasks does the FIP algorithm target and how does it perform?",{"text":84,"@type":76},"FIP targets continual learning and network sparsification objectives and reports performance comparable to state of the art on language models (e.g., BERT), vision transformers, and CNNs with modest computation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]