[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117609-en":3,"doc-seo-117609-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117609,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Refined Convergence and Topology Learning for Decentralized SGD with Heterogeneous Data","Decentralized and federated learning require algorithms that remain efficient under strongly heterogeneous data distributions across agents. This paper revisits Decentralized Stochastic Gradient Descent (D-SGD) and identifies neighborhood heterogeneity as the key quantity governing the convergence rate. By coupling communication topology with heterogeneity, the analysis clarifies how these factors interact in convergence time. A data-dependent topology learning criterion is proposed to reduce, and potentially eliminate, the harmful impact of heterogeneity. For label-skew classification, the topology-learning task is formulated as a tractable optimization problem solved via a Frank–Wolfe method, producing sparse topologies that balance convergence speed and per-iteration communication cost, validated on simulated and real-world experiments.","Reﬁned Convergence and Topology Learning for Decentralized SGD with Heterogeneous Data  \nBatiste Le Bars Aurélien Bellet Marc Tommasi  \nUniv. Lille, Inria, CNRS, Centrale Lille, UMR 9189, CRIStAL, F-59000 Lille  \nErick Lavoie Anne-Marie Kermarrec  \nUniversité de Bâle, Bâle, Switzerland EPFL, Lausanne, Switzerland  \nAbstract  \nOne of the key challenges in decentralized and federated learning is to design algorithms that efﬁciently deal with highly heterogeneous data distributions across agents. In this paper, we revisit the analysis of the popular Decentralized Stochastic Gradient Descent algorithm (D-SGD) under data heterogeneity. We exhibit the key role played by a new quantity, called neighborhood heterogeneity, on the convergence rate of D-SGD.  \nBy coupling the communication topology and the heterogeneity, our analysis sheds light on the poorly understood interplay between these two concepts. We then argue that neighborhood heterogeneity provides a natural criterion to learn data-dependent topologies that reduce (and can even eliminate) the otherwise detrimental effect of data heterogeneity on the convergence time of D-SGD. For the important case of classiﬁcation with label skew, we formulate the problem of learning such a good topology as a tractable optimization problem that we solve with a FrankWolfe algorithm. As illustrated over a set of simulated and real-world experiments, our approach provides a principled way to design a sparse topology that balances the convergence speed and the per-iteration communication costs of D-SGD under data heterogeneity.  \n1 INTRODUCTION  \nDecentralized and federated learning methods allow training from data stored locally by several agents (nodes) without  \nProceedings of the 26th International Conference on Artiﬁcial Intelligence and Statistics (AISTATS) 2023, Valencia, Spain. PMLR: Volume 206 . Copyright 2023 by the author(s) .  \nexchanging raw data, in line with the increasing demand for more privacy-preserving algorithms (Kairouz et al., 2021) . One of the key challenges in decentralized learning is to deal with data heterogeneity: as each agent collects its own data, local datasets typically exhibit different distributions. In this work, we study this challenge in the context of fully decentralized learning algorithms, which provide a scalable and robust alternative to server-based approaches (Colin et al., 2016 ; Lian et al., 2017 ; Koloskova et al., 2019, 2020) . Fully decentralized optimization algorithms, such as the celebrated Decentralized SGD (D-SGD) (Lian et al., 2017, 2018 ; Koloskova et al., 2020), operate on a graph representing the communication topology, i.e. which pairs of nodes exchange information with each other. The connectivity of the topology then rules a trade-off between the convergence rate and the per-iteration communication complexity of fully decentralized algorithms (Wang et al., 2019) . Choosing a good topology for fully decentralized machine learning is therefore an important question, and remains a largely open problem in the presence of data heterogeneity.  \nUntil recently, the impact of the communication topology on the convergence was believed to be mainly characterized by its spectral gap: a large spectral gap indicating good connectivity and thus faster convergence. Focusing solely on the connectivity of the topology has however shown tobe insufﬁcient, even when we have identically distributed data (Neglia et al., 2020 ; Vogels et al., 2022) . In the heterogeneous setting, Bellet et al. (2022) notably observe that the choice of topology has a large inﬂuence, beyond its spectral gap, on the convergence speed of D-SGD. However, these empirical observations are not supported by any theory.  \nIn this work, we ﬁll the theoretical gap that currently existson these questions. We focus on D-SGD (Lian et al., 2017, 2018 ; Koloskova et al., 2020), which is arguably the most popular decentralized optimization algorithm in the context of machine learning due ","cbCaiaJBXI5oUZKM","https://ap.wps.com/l/cbCaiaJBXI5oUZKM","pdf",1470866,1,31,"English","en",105,"# Introduction\n## Decentralized and federated learning under data heterogeneity\n## Topology and convergence: spectral gap limits\n## Refined analysis via neighborhood heterogeneity\n## Learning data-dependent topologies and Frank-Wolfe","[{\"question\":\"What is the main challenge addressed in decentralized learning?\",\"answer\":\"Designing decentralized optimization algorithms that efficiently handle highly heterogeneous data distributions across agents without exchanging raw data.\"},{\"question\":\"How does neighborhood heterogeneity affect D-SGD convergence?\",\"answer\":\"Neighborhood heterogeneity, which couples communication topology with local data distributions, determines the convergence rate beyond topology connectivity or spectral gap alone.\"},{\"question\":\"How is a good data-dependent topology learned in the label-skew setting?\",\"answer\":\"The paper formulates topology learning as a tractable optimization problem and solves it with a Frank–Wolfe algorithm to construct a sparse topology that balances convergence speed and communication cost.\"}]","Refined Convergence and Topology Learning for Decentralized SGD with Heterogeneous Data | PDF",1785677266,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"refined-convergence-and-topology-learning-for-decentralized-sgd-with-heterogeneous-data","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/refined-convergence-and-topology-learning-for-decentralized-sgd-with-heterogeneous-data/117609/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main challenge addressed in decentralized learning?","Question",{"text":76,"@type":77},"Designing decentralized optimization algorithms that efficiently handle highly heterogeneous data distributions across agents without exchanging raw data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does neighborhood heterogeneity affect D-SGD convergence?",{"text":81,"@type":77},"Neighborhood heterogeneity, which couples communication topology with local data distributions, determines the convergence rate beyond topology connectivity or spectral gap alone.",{"name":83,"@type":74,"acceptedAnswer":84},"How is a good data-dependent topology learned in the label-skew setting?",{"text":85,"@type":77},"The paper formulates topology learning as a tractable optimization problem and solves it with a Frank–Wolfe algorithm to construct a sparse topology that balances convergence speed and communication cost.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]