[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85112-en":3,"doc-seo-85112-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85112,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","SLORR: Simple and Efficient In-Training Low-Rank Regularization","Low-rank factorization compresses neural networks, yet aggressive factorization is often inaccurate for modern architectures. Training-time low-rank regularizers improve compressibility but typically rely on costly SVDs, alter model architecture via factorized layers, or depend on cached spectral quantities. SLORR introduces a stateless, architecture-preserving framework for in-training low-rank regularization using GPU-friendly approximations. Two variants target the Hoyer sparsity metric and the nuclear norm, with approximation guarantees for values and gradients.","arXiv :2607 .08754v 1 [ cs .LG] 9 Jul 2026  \nSLORR: Simple and Efficient In-Training Low-Rank  \nRegularization  \nDavid González-Martínez 1,2,3,4 Shiwei Liu 1,3,4  \n1Max Planck Institute for Intelligent Systems 2University of Tübingen  \n3ELLIS Institute Tübingen 4Tübingen AI Center  \n[david.martinez@tuebingen.mpg.de](david.martinez@tuebingen.mpg.de) [sliu@tue.ellis.eu](sliu@tue.ellis.eu)  \nAbstract  \nLow-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. Existing training-time low-rank regularizers can improve compressibility, but they often require SVDs of large weight matrices, modify the model architecture (introducing additional trainable parameters), or rely on stateful cached quantities. To address these limitations, we introduce SLORR, a simple, stateless, and architecture-preserving framework for in-training low-rank regularization, instantiated with two main variants based on the Hoyer sparsity metric and the nuclear norm. SLORR directly regularizes the original weight matrices using GPU-friendly approximations for the forward and backward passes of the regularizers, for which we provide approximation guarantees. We first evaluate SLORR on ImageNet-1K across short-horizon continued training of ResNet-50, ViT-B/16, and ViT-L/16, and pretraining of ResNet-18, where SLORR induces compressibility while introducing less than 8% training overhead. We further evaluate SLORR-Hoyer in LLM pretraining at 135M and 560M scales: SLORR-trained compressed models preserve performance substantially better than unregularized models while adding less than 1% average training overhead.  \n1 Introduction  \nLow-rank factorization is a widely used approach for reducing the computational and memory costs of neural networks: a dense weight matrix is replaced by a product of smaller factors, yielding cheaper inference when the retained rank is sufficiently small [1–4] . However, compression quality depends on the spectral properties of weight matrices: discarding meaningful spectral directions can degrade model performance. Thus, while moderate post-training compression is generally possible, modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. This motivates training-time methods that encourage low-rank structure before compression. A natural formulation is to encourage sparsity in the singular values of weight matrices, concentrating spectra into fewer dominant directions. Although direct spectral regularization can be effective [5, 6], large singular value decompositions (SVDs) are needed at every training iteration, which is prohibitively expensive for modern training pipelines.  \nSeveral ways of avoiding SVDs have been proposed [7–9]; however, these generally introduce new limitations that hinder practicality. For example, some works replace layers with factorized parameterizations, changing the architecture [7, 8], but this generally increases the number of trainable parameters and alters optimization dynamics. A further class of methods maintains cached spectral quantities that are only periodically updated [9], reducing but not eliminating SVD cost, and introducing statefulness and additional hyperparameters.  \nPreprint.  \nTable 1: Comparison of different low-rank-inducing regularizers. Detailed discussion in Section 2. Method Changes arch. SVD Efficiency & scalability Stateful Prior target rank  \nSVD-based [5] Factorize + reg. [7] LoRITa [8] Q3R [9]  \nSLORR (ours)  \nNo  \nYes  \nYes  \nNo  \nNo  \nYes  \nNo  \nNo  \nPeriodic  \nNo  \nPoor  \nGood  \nGood  \nModerate  \nGood  \nNo  \nNo  \nNo  \nYes  \nNo  \nNo  \nNo  \nNo  \nYes  \nNo  \nOur proposed method, SLORR, addresses these limitations. In particular, SLORR operates directly on the original weight matrices without altering the architecture, is SVD-free and efficient, and does not maintain any cached quantities, providing a v","cbCaird50f69TuBs","https://ap.wps.com/l/cbCaird50f69TuBs","pdf",1261783,5,1,41,"English","en",105,"# Abstract\n# Introduction\n## Low-rank factorization and training-time challenges\n## Avoiding SVDs and practical limitations\n# Background and Related Work","[{\"question\":\"What problem does SLORR address in low-rank training regularization?\",\"answer\":\"SLORR targets the gap between commonly used low-rank factorization and the difficulty of achieving it during training without accuracy loss. It addresses shortcomings of existing regularizers that are expensive, require architecture changes, or depend on cached quantities.\"},{\"question\":\"How does SLORR avoid expensive SVD computations during training?\",\"answer\":\"SLORR is SVD-free and stateless, approximating the spectral quantities needed for regularizer forward and backward passes using GPU-efficient polar factor approximations.\"},{\"question\":\"What variants of SLORR are proposed, and what do they regularize?\",\"answer\":\"SLORR has two main variants: SLORR-Hoyer, corresponding to the squared Hoyer sparsity metric, and SLORR-Nuc, corresponding to the nuclear norm. Both directly regularize the original weight matrices while preserving gradients with approximation guarantees.\"}]",1784201175,103,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"slorr-simple-and-efficient-in-training-low-rank-regularization","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/slorr-simple-and-efficient-in-training-low-rank-regularization/85112/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does SLORR address in low-rank training regularization?","Question",{"text":76,"@type":77},"SLORR targets the gap between commonly used low-rank factorization and the difficulty of achieving it during training without accuracy loss. It addresses shortcomings of existing regularizers that are expensive, require architecture changes, or depend on cached quantities.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SLORR avoid expensive SVD computations during training?",{"text":81,"@type":77},"SLORR is SVD-free and stateless, approximating the spectral quantities needed for regularizer forward and backward passes using GPU-efficient polar factor approximations.",{"name":83,"@type":74,"acceptedAnswer":84},"What variants of SLORR are proposed, and what do they regularize?",{"text":85,"@type":77},"SLORR has two main variants: SLORR-Hoyer, corresponding to the squared Hoyer sparsity metric, and SLORR-Nuc, corresponding to the nuclear norm. Both directly regularize the original weight matrices while preserving gradients with approximation guarantees.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]