[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118706-en":3,"doc-seo-118706-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118706,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication - research summary","Decentralized stochastic optimization is studied on a fixed communication graph where each of n machines holds part of the objective and communicates only with neighbors. To mitigate communication limits, updates are compressed via quantization or sparsification using both unbiased and biased compression operators, characterized by a quality parameter where no compression corresponds to 􀀎=1. A gossip-based stochastic gradient descent method, CHOCO-SGD, is proposed for strongly convex objectives with a convergence rate depending on n, iterations, eigengap, and compression quality, while CHOCO-GOSSIP attains linear-time consensus performance. Experiments show improved baselines and large reductions in communication.","Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication  \nAnastasia Koloskova * 1 Sebastian U. Stich * 1 Martin Jaggi 1  \nAbstract  \nWe consider decentralized stochastic optimization with the objective function (e.g. data samples for machine learning tasks) being distributed over n machines that can only communicate to their neighbors on a ﬁxed communication graph. To address the communication bottleneck, the nodes compress (e.g. quantize or sparsify) their model updates. We cover both unbiased and biased compression operators with quality denoted by 􀀎 􀀔 1 (􀀎 = 1 meaning no compression) .  \nWe (i) propose a novel gossip-based stochastic gradient descent algorithm, CHOCO-SGD, that converges at rate O 􀀀1=(nT ) + 1=(T 􀀚2 􀀎)2 􀀁 for strongly convex objectives, where T denotes the number of iterations and 􀀚 the eigengap of the connectivity matrix. We (ii) present a novel gossip algorithm, CHOCO-GOSSIP, for the average consensus problem that converges in time O(1=(􀀚2 􀀎) log(1=\")) for accuracy \" > 0. This is (up to our knowledge) the ﬁrst gossip algorithm that supports arbitrary compressed messages for  \n􀀎 > 0 and still exhibits linear convergence. We (iii) show in experiments that both of our algorithms do outperform the respective state-of-theart baselines and CHOCO-SGD can reduce communication by at least two orders of magnitudes.  \n1. Introduction  \nDecentralized machine learning methods are becoming core aspects of many important applications, both in view ofscalability to larger datasets and systems, but also from the perspective of data locality, ownership and privacy. We consider decentralized optimization methods that do not rely on a central coordinator (e.g. parameter server) but instead only require on-device computation and local communica-  \n*Equal contribution 1EPFL, Lausanne, Switzerland. Correspondence to: Anastasia Koloskova \u003Canastasia.koloskova@epﬂ.ch> .  \nProceedings of the 36 th International Conference on Machine Learning, Long Beach, California, PMLR 97, 2019 . Copyright 2019 by the author(s) .  \ntion with neighboring devices. This covers for instance the classic setting of training machine learning models in large data-centers, but also emerging applications were the computations are executed directly on the consumer devices, which keep their part of the data private at all times.1  \nFormally, we consider optimization problems distributed across n devices or nodes of the form  \nf? := mx2inRd 􀀔f (x) := 1n Xi1 fi (x)􀀕 ; (1)  \nwhere fi : Rd ! R for i 2 [n] := f1; : : : ; ng are the objectives deﬁned by the data available locally on each node. We also allow each local objective fi to have stochastic optimization (or sum) structure, covering the important case of empirical risk minimization in distributed machine learning and deep learning applications.  \nDecentralized Communication. We model the network topology as a graph where edges represent the communication links along which messages (e.g. model updates) can be exchanged. The decentralized setting is motivated by centralized topologies (corresponding to a star graph) often not being possible, and otherwise often posing a signiﬁcant bottleneck on the central node in terms of communication latency, bandwidth and fault tolerance. Decentralized topologies avoid these bottlenecks and thereby offer hugely improved potential in scalability. For example, while the master node in the centralized setting receives (and sends) in each round messages from all workers, 􀀂(n) in total2 , in decentralized topologies the maximal degree of the network is often constant (e.g. ring or torus) or a slowly growing function in n (e.g. scale-free networks) .  \nDecentralized Optimization. For the case of deterministic (full-gradient) optimization, recent seminal theoretical advances show that the network topology only affects higher-order terms of the convergence rate of decentralized optimization algorithms on convex problems (Scaman et al.,  \n1Note th","cbCaiiXMKMTY0xAP","https://ap.wps.com/l/cbCaiiXMKMTY0xAP","pdf",4981541,1,24,"English","en",105,"# Abstract\n# Introduction\n## Decentralized Communication\n## Decentralized Optimization\n## Communication Compression\n## Contributions","[{\"question\":\"What problem setting is considered for the decentralized optimization task?\",\"answer\":\"The objective is distributed across n machines on a fixed communication graph, and each node can communicate only with its neighbors. The study covers decentralized stochastic gradient descent where local objectives may have stochastic optimization structure.\"},{\"question\":\"How does the work reduce communication overhead in distributed training?\",\"answer\":\"Nodes compress model updates using operators such as quantization or sparsification. The analysis includes both unbiased and biased compression with a quality factor that controls how accurately compressed updates represent the originals.\"},{\"question\":\"What are CHOCO-SGD and CHOCO-GOSSIP designed to achieve?\",\"answer\":\"CHOCO-SGD targets decentralized stochastic gradient descent on strongly convex objectives, while CHOCO-GOSSIP solves the average consensus problem. Both methods incorporate gossip-style communication and provable convergence guarantees under compressed messages.\"}]","Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication - research summary | PDF",1785685005,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"decentralized-stochastic-optimization-and-gossip-algorithms-with-compressed-communication-research-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/decentralized-stochastic-optimization-and-gossip-algorithms-with-compressed-communication-research-summary/118706/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem setting is considered for the decentralized optimization task?","Question",{"text":75,"@type":76},"The objective is distributed across n machines on a fixed communication graph, and each node can communicate only with its neighbors. The study covers decentralized stochastic gradient descent where local objectives may have stochastic optimization structure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the work reduce communication overhead in distributed training?",{"text":80,"@type":76},"Nodes compress model updates using operators such as quantization or sparsification. The analysis includes both unbiased and biased compression with a quality factor that controls how accurately compressed updates represent the originals.",{"name":82,"@type":73,"acceptedAnswer":83},"What are CHOCO-SGD and CHOCO-GOSSIP designed to achieve?",{"text":84,"@type":76},"CHOCO-SGD targets decentralized stochastic gradient descent on strongly convex objectives, while CHOCO-GOSSIP solves the average consensus problem. Both methods incorporate gossip-style communication and provable convergence guarantees under compressed messages.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]