[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124954-en":3,"doc-seo-124954-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124954,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Dependable Distributed Training of Compressed Machine Learning Models","Distributed training and model compression have improved machine learning efficiency, yet existing approaches largely ignore how learning quality is distributed across outcomes, relying instead on average performance. This omission reduces the dependability of trained models, whose realized accuracy or loss may fall far below expectations. The work introduces DepL, a dependable learning orchestration framework that selects data sources, model variants, and node clusters to achieve a target learning quality with a specified probability while minimizing training cost. The paper proves constant-competitive performance and polynomial complexity and shows over 27% gains versus state of the art.","POLITECNICO DI TORINO Repository ISTITUZIONALE  \nDependable Distributed Training of Compressed Machine Learning Models  \nOriginal  \nDependable Distributed Training of Compressed Machine Learning Models / Malandrino, F. ; Di Giacomo, G. ; Levorato, M. ; Chiasserini, C. F.. -ELETTRONICO. - (2024) . (Intervento presentato al convegno IEEE WoWMoM 2024 tenutosi a Perth (Australia) nel 04-07 June 2024) [10 . 1109/WoWMoM60985 .2024.00036] .  \nAvailability:  \nThis version is available at: 11583/2986252 since: 2024-02-22T16:49:38Z  \nPublisher: IEEE  \nPublished  \nDOI:10.1109/WoWMoM60985.2024.00036  \nTerms of use:  \nThis article is made available under terms and conditions as specified in the corresponding bibliographic description in the repository  \nPublisher copyright  \nIEEE postprint/Author's Accepted Manuscript  \n©2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collecting works, for resale or lists, or reuse of any copyrighted component of this work in other works.  \n(Article begins on next page)  \n18 September 2024  \nDependable Distributed Training of Compressed Machine Learning Models  \nFrancesco Malandrino∗†, Giuseppe Di Giacomo‡, Marco Levorato§ , Carla Fabiana Chiasserini‡∗†  \n∗ CNR-IEIIT, Italy – †CNIT, Italy – ‡Politecnico di Torino, Italy – § UC Irvine, USA  \nAbstract—The existing work on the distributed training of machine learning (ML) models has consistently overlooked the distribution of the achieved learning quality, focusing instead on its average value. This leads to a poor dependability of the resulting ML models, whose performance may be much worse than expected. We fill this gap by proposing DepL, a framework for dependable learning orchestration, able to make high-quality, efficient decisions on (i) the data to leverage for learning, (ii) the models to use and when to switch among them, and (iii) the clusters of nodes, and the resources thereof, to exploit. For concreteness, we consider as possible available models a full DNN and its compressed versions. Unlike previous studies, DepL guarantees that a target learning quality is reached with a target probability, while keeping the training cost at a minimum. We prove that DepL has constant competitive ratio and polynomial complexity, and show that it outperforms the state-of-the-art by over 27% and closely matches the optimum.  \nI. INTRODUCTION  \nMachine learning (ML) models, and deep neural networks (DNNs) in particular, are becoming more capable, but also harder and more costly to train. This issue has been tackled through two complementary approaches: distributed training and model compression. The former exploits the resources (e.g., computational) and data available at different nodes, thus better distributing the training burden. The latter includes a wealth of different techniques (e.g., model pruning and knowledge distillation) allowing for the transfer of knowledge from a model to a (usually, simpler) one.  \nDistributed training and model compression have been variously combined together in the literature, allowing the used resources to adapt to the chosen model [1], selecting the best model given the available resources [2], [3], and even performing mutual adaptation between the two [4] . However, an aspect that has been overlooked so far, which we aim to investigate in this work, is the dependability of DNN training. Indeed, ML is increasingly used for mission-critical (and even safety-related) applications, hence traditional approaches focusing on the expected learning quality may be inadequate to present needs.  \nIn this work, we aim at providing the right network support to make DNN training dependable, i.e., to guarantee that a target learning quality (e.g., loss or accuracy) is reached by a target time and with a target probability – at the minimum cost.","cbCairQkXNsKiLuW","https://ap.wps.com/l/cbCairQkXNsKiLuW","pdf",2129344,1,11,"English","en",105,"# Introduction\n## Motivation and problem setting\n## Key observation and dependability factors\n# Proposed framework\n## DepL orchestration decisions\n## Guarantees and performance analysis\n# Experimental evaluation\n## Comparison with state of the art","[{\"question\":\"Why does focusing on average learning quality hurt the dependability of distributed ML models?\",\"answer\":\"Because performance can vary widely across training outcomes, models may reach much worse quality than expected even if the average looks good. This makes the resulting ML system less reliable for mission-critical use.\"},{\"question\":\"What decisions does DepL make during dependable distributed training?\",\"answer\":\"DepL jointly decides which data sources to use, which model variants to train (e.g., full DNN or compressed versions) and when to switch between them, and which clusters of nodes and resources to exploit at different stages.\"},{\"question\":\"How does DepL ensure a target learning quality with a target probability while controlling cost?\",\"answer\":\"DepL is designed to guarantee that the target quality (such as loss or accuracy) is reached with a specified probability, while minimizing the training cost. The paper also establishes constant competitive ratio and polynomial complexity.\"}]","Dependable Distributed Training of Compressed Machine Learning Models | PDF",1785895590,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dependable-distributed-training-of-compressed-machine-learning-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dependable-distributed-training-of-compressed-machine-learning-models/124954/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does focusing on average learning quality hurt the dependability of distributed ML models?","Question",{"text":75,"@type":76},"Because performance can vary widely across training outcomes, models may reach much worse quality than expected even if the average looks good. This makes the resulting ML system less reliable for mission-critical use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What decisions does DepL make during dependable distributed training?",{"text":80,"@type":76},"DepL jointly decides which data sources to use, which model variants to train (e.g., full DNN or compressed versions) and when to switch between them, and which clusters of nodes and resources to exploit at different stages.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DepL ensure a target learning quality with a target probability while controlling cost?",{"text":84,"@type":76},"DepL is designed to guarantee that the target quality (such as loss or accuracy) is reached with a specified probability, while minimizing the training cost. The paper also establishes constant competitive ratio and polynomial complexity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]