[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82492-en":3,"doc-seo-82492-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82492,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Entropy Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning","Federated Learning enables multiple clients to train collaboratively without sharing data, yet it is hindered by data heterogeneity and partial client participation. Learning sparse models improves communication and computation, but small-sample high-dimensional settings (d≫N) make optimization unstable and can yield poor generalization. The work proposes probabilistic gates with an L0 constraint, enhanced by entropy regularization, to prevent premature commitment to sparse support. Experiments on synthetic and real benchmarks improve test accuracy and sparsity recovery over Fed-IHT and pruning-after-FedAvg under multiple heterogeneities.","arXiv :2607 .00275v1 [ cs .LG] 30 Jun 2026  \nEntropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning  \nKrishna Harsha Kovelakuntla Huthasana  \nAlireza Olama Andreas Lundell  \nDepartment of Engineering and Information Technology, Åbo Akademi University {kkovelak, alireza . olama, [andreas. lundell}@abo. fi](andreas. lundell}@abo. fi)  \nJuly 2, 2026  \nAbstract  \nFederated Learning (FL) is a distributed machine learning (ML) paradigm with collaboration among multiple clients without sharing data. FL is challenging under data heterogeneity and partial client participation. Learning sparse models is useful for communication and computational efficiency in FL, but it is especially difficult in the small-sample high-dimensional regime (d ≫ N ) where optimization can yield parameter configurations that fail to generalize to unseen test data. While magnitude-based pruning doesn’t account for uncertainty exploration in the parameter space, a formulation with probabilistic gates and an L0 constraint allows sampling from competing sparse configurations during training. In this work, we study entropy regularization of gate distributions as a mechanism to maintain uncertainty in sparse federated optimization by preventing early commitment to sparse support. We examine its impact under data heterogeneity, client participation heterogeneity, and sparsity. Experiments on synthetic and real-world benchmarks show consistent improvements over federated iterative hard thresholding (Fed-IHT) and pruning after dense federated averaging (FedAvg) training, both in statistical performance on test data and in sparsity recovery accuracy.  \nKeywords  \nEntropy Regularization , Sparsity , Federated Learning , Uncertainty , Parameter Exploration , Probabilistic Gates , Entropy Maximization , Sparse Federated Learning , L0 Constraint  \n1 Introduction  \nFederated Learning (FL) algorithms operate in a distributed machine learning (ML) setting in which multiple clients collaborate to train [16] . This framework is characterized by privacy requirements of each client and avoids data sharing. While not all distributed settings are privacy-sensitive, FL is still beneficial because it eliminates the need for centralized data [11] . FL can be coordinated either by a single server or by clients communicating with one another. In this work, we study FL with central orchestration by a server to obtain a single global model as shown in Figure 1. A global model is typically learned by iteratively averaging parameters or gradients from clients and redistributing the global model to clients for further learning. However, statistical heterogeneity across clients and partial client participation during training pose challenges to the learning process in FL. Furthermore, sparse training and inference are desirable to improve generalizability [23] and enhance computational and communication efficiency in FL, thereby posing an additional challenge of discovering sparse models [25] .  \nA common approach to inducing sparsity relies on L1 and L2 norms [1] for regularization, which depend directly on parameter magnitudes and offer varying levels of shrinkage. In contrast, using a magnitudeindependent L0 pseudo-norm is advantageous because it imposes a constant penalty on nonzero parameters and is useful for learning a model with a desired parameter density ρ . The Lagrangian for the L 0 density-constrained optimization problem in FL can be defined as:  \nC |θ|  \nL (θ, λ) =X n~~ ~~cN L (c)(θ) + λ (∥θ∥0 − ρ|θ|) , ∥θ∥0 = X I [θj  0], (1)  \nc=1 j=1  \n[1]For θ ∈ Rd , L1 norm is ∥θ∥1 = Pdi=1 |θ i| and L2 norm is ∥θ∥2 = 􀀐 Pdi=1 θ2i􀀑 1/2 .  \nClient 1 Client 2 Client 3 Client 4  \nFigure 1: Client–server federated learning architecture with central orchestration. Solid arrows indicate the aggregated global model distributed to clients, while dotted arrows indicate local model updates sent from clients to the server.  \nwhere, L(c)(θ) denotes the norm","cbCaihF0dFcdjDP0","https://ap.wps.com/l/cbCaihF0dFcdjDP0","pdf",4007145,2,1,11,"English","en",105,"# Abstract\n# Keywords\n# 1 Introduction\n## Federated Learning setup and challenges\n## Sparsity via L1/L2 regularization vs L0 density constraint\n## Entropy-regularized probabilistic gates background and prior work","[{\"question\":\"Why is sparse model discovery difficult in scarce-data federated learning?\",\"answer\":\"In small-sample high-dimensional regimes (d≫N), optimization becomes unstable and can converge to parameter configurations that recover sparsity poorly and generalize weakly to unseen test data.\"},{\"question\":\"What role do probabilistic gates and the L0 constraint play?\",\"answer\":\"Probabilistic gates introduce stochastic sampling of competing sparse configurations during training, while an L0 constraint targets a desired parameter density and avoids dependence on parameter magnitudes alone.\"},{\"question\":\"How does entropy regularization affect training with probabilistic gates?\",\"answer\":\"Entropy regularization maintains uncertainty in the gate distribution, helping prevent early commitment to sparse support and improving sparse federated optimization behavior.\"}]",1784180897,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"entropy-regularized-probabilistic-gates-for-sparse-model-discovery-in-scarce-data-federated-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/entropy-regularized-probabilistic-gates-for-sparse-model-discovery-in-scarce-data-federated-learning/82492/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is sparse model discovery difficult in scarce-data federated learning?","Question",{"text":75,"@type":76},"In small-sample high-dimensional regimes (d≫N), optimization becomes unstable and can converge to parameter configurations that recover sparsity poorly and generalize weakly to unseen test data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What role do probabilistic gates and the L0 constraint play?",{"text":80,"@type":76},"Probabilistic gates introduce stochastic sampling of competing sparse configurations during training, while an L0 constraint targets a desired parameter density and avoids dependence on parameter magnitudes alone.",{"name":82,"@type":73,"acceptedAnswer":83},"How does entropy regularization affect training with probabilistic gates?",{"text":84,"@type":76},"Entropy regularization maintains uncertainty in the gate distribution, helping prevent early commitment to sparse support and improving sparse federated optimization behavior.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]