[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120643-en":3,"doc-seo-120643-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120643,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","The Hidden Vulnerability of Distributed Learning in Byzantium","While machine learning achieves celebrated success, distributed SGD remains vulnerable to Byzantine workers that inject poisoned gradients during training. Prior work proposed Byzantine-resilient aggregation methods that guarantee convergence by ensuring SGD reaches a minimization target despite a minority of adversaries. This paper shows convergence alone is insufficient in high-dimensional non-convex settings: an attacker exploits loss non-convexity to force convergence toward ineffective models, with poisoning margins growing at least like √p(d). A simple attack is demonstrated on CIFAR-10 and MNIST, and the Bulyan method is introduced to significantly reduce the attacker's leeway to a tight O(1/√p(d)) bound.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nThe Hidden Vulnerability of Distributed Learning in Byzantium  \nEl Mahdi El Mhamdi 1 Rachid Guerraoui 1 Sbastien Rouault 1  \nAbstract  \nWhile machine learning is going through an era of celebrated success, concerns have been raised about the vulnerability of its backbone: stochastic gradient descent (SGD) . Recent approaches have been proposed to ensure the robustness of distributed SGD against adversarial (Byzantine) workers sending poisoned gradients during the training phase. Some of these approaches have been proven Byzantine–resilient: they ensure the convergence of SGD despite the presence of a minority of adversarial workers. We show in this paper that convergence is not enough. In high dimension d 􀀝 1, an adversary can build on the loss function's non–convexity to make SGD converge to ineffective models. More precisely, we bring to light that existing Byzantine–resilient schemes leave a margin of poisoning of 􀀊(f(d)), where f (d) increases at least like pd. Based on this leeway, we build a simple attack, and experimentally show its strong to utmost effectivity on CIFAR–10 and MNIST. We introduce Bulyan, and prove it signiﬁcantly reduces the attacker's leeway to a narrow O (1=pd ) bound. We empirically show that Bulyan does not suffer the fragility of existing aggregation rules and, at a reasonable cost in terms of required batch size, achieves convergence as if only non–Byzantine gradients had been used to update the model.  \n1. Introduction  \nStochastic Gradient Descent (SGD), is arguably the backbone of the most successful machine learning methods (LeCun et al., 2015; Abadi et al., 2016; Dean et al., 2012; Bottou, 1998) . Gradient Descent (GD), the underlying principle of SGD, is so straightforward that, as sometimes said,“Newton could have invented it in his time”. In particular, GD relies on a simple observation: given a function  \n1EPFL, Lausanne, Switzerland. Correspondence to: (without spaces) \u003Cﬁrstname.lastname@epﬂ.ch> .  \nProceedings of the 35 th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018 . Copyright 2018 by the author(s) .  \nQ, depending on a parameter x, if one keeps updating x in the opposite direction of the gradient of Q, with reasonably small, but not too small (Bottou, 1998) steps, x eventually reaches the global minimum if Q is convex or, if Q is not convex 1 , reach a region where Q is either ﬂat or in some local minima. SGD is the lightweight version of GD, where a sample is drawn at random to estimate the gradient of Q.  \nBeyond image recognition or video labeling for social networks, SGD–based machine learning is venturing into safety–critical applications, like health–care (Holzinger, 2016) and transportation (Bojarski et al., 2016) . Meanwhile, a growing body of work, coined Adversarial Machine Learning (Biggio & Roli, 2017; Goodfellow et al., 2014; Gilmer et al., 2018; Kumar et al., 2017) is unveiling serious vulnerabilities in some of the most performing algorithms. Essentially, the general effort towards robust ML is conducted against three kinds of attacks: poisoning ones (Biggio & Laskov, 2012; Koh & Liang, 2017) where an adversary injects poisoned data during the training phase, exploratory attacks, where a curious attacker attempts to infer privacy–sensitive information, and evasion attacks, where attackers try to fool an already trained model with adversarial inputs. The three fronts are complementary and each kind of attack poses a challenge on its own.  \nIn the context of poisoning attacks, an emerging line of research looks at robustness through the lenses of (distributed) optimization (Chen et al., 2017; Su, 2017; Blanchard et al., 2017) . Interestingly, SGD can be proven to converge despite the presence of a (bounded) number of adversaries. The general r","cbCaigGMYaZhTrcX","https://ap.wps.com/l/cbCaigGMYaZhTrcX","pdf",822246,1,13,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why is convergence of distributed SGD not sufficient in this work?\",\"answer\":\"Because in high-dimensional non-convex neural-network loss landscapes, an adversary can exploit non-convexity so that SGD converges to ineffective models, even if convergence is still guaranteed.\"},{\"question\":\"What capability does a Byzantine adversary have according to the paper?\",\"answer\":\"The adversary can craft poisoned gradients to take advantage of the loss function’s non-convexity, creating a poisoning margin large enough to steer the training outcome toward poor models.\"},{\"question\":\"How does Bulyan change the attacker’s impact?\",\"answer\":\"The paper introduces Bulyan and proves it significantly reduces the attacker’s leeway to a narrow O(1/√p(d)) bound, while empirically avoiding fragility seen in existing aggregation rules and enabling convergence comparable to using only non-Byzantine gradients.\"}]","The Hidden Vulnerability of Distributed Learning in Byzantium | PDF",1785731058,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-hidden-vulnerability-of-distributed-learning-in-byzantium","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-hidden-vulnerability-of-distributed-learning-in-byzantium/120643/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is convergence of distributed SGD not sufficient in this work?","Question",{"text":75,"@type":76},"Because in high-dimensional non-convex neural-network loss landscapes, an adversary can exploit non-convexity so that SGD converges to ineffective models, even if convergence is still guaranteed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What capability does a Byzantine adversary have according to the paper?",{"text":80,"@type":76},"The adversary can craft poisoned gradients to take advantage of the loss function’s non-convexity, creating a poisoning margin large enough to steer the training outcome toward poor models.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Bulyan change the attacker’s impact?",{"text":84,"@type":76},"The paper introduces Bulyan and proves it significantly reduces the attacker’s leeway to a narrow O(1/√p(d)) bound, while empirically avoiding fragility seen in existing aggregation rules and enabling convergence comparable to using only non-Byzantine gradients.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]