[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125530-en":3,"doc-seo-125530-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125530,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Towards Inference Delivery Networks - Distributing Machine Learning with Optimality Guarantees","Machine learning inference increasingly powers real-time applications, but two common deployment choices—running models locally on devices or offloading to remote cloud services—often fail to jointly satisfy accuracy and delay constraints. The paper introduces inference delivery networks (IDNs), coordinating heterogeneous computing nodes across access, edge, regional data centers, and cloud to deliver the best latency–accuracy trade-off. It proposes a distributed dynamic policy that allocates and updates model sets using recent request observations and limited neighbor communication, providing strong adversarial performance guarantees and outperforming greedy heuristics in realistic scenarios.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \n[provided by](provided by arXiv.org)[ arXiv.org](provided by arXiv.org) e-Print Archive  \nTowards Inference Delivery Networks: Distributing Machine Learning with Optimality Guarantees  \nTareq Si Salem􀀃 , Gabriele Castellano􀀃y , Giovanni Neglia􀀃 , Fabio Pianesey , Andrea Araldoz 􀀃 Inria, Universit Cte d'Azur, France, [f](ftareq.si-salem)[tareq.si-salem](ftareq.si-salem), gabriele.castellano, [giovanni.neglia](giovanni.negliag@inria.fr)[g](giovanni.negliag@inria.fr)[@inria.fr](giovanni.negliag@inria.fr), y Nokia Bell Labs, France, ffabio.pianese, [gabriele.castellano.ext](gabriele.castellano.extg@nokia.com)[g](gabriele.castellano.extg@nokia.com)[@nokia.com](gabriele.castellano.extg@nokia.com), z Tlcom SudParis -Institut Polytechnique de Paris, France, [andrea.araldo@telecom-sudparis.eu](andrea.araldo@telecom-sudparis.eu)  \narXiv :2105 .025 10v2 [ cs .NI] 24 Jul 2021  \nAbstract—An increasing number of applications rely on complex inference tasks that are based on machine learning (ML). Currently, there are two options to run such tasks: either they are served directly by the end device (e.g., smartphones, IoT equipment, smart vehicles), or ofﬂoaded to a remote cloud. Both options may be unsatisfactory for many applications: local models may have inadequate accuracy, while the cloud may fail to meet delay constraints. In this paper, we present the novel idea of inference delivery networks (IDNs), networks of computing nodes that coordinate to satisfy ML inference requests achieving the best trade-off between latency and accuracy. IDNs bridge the dichotomy between device and cloud execution by integrating inference delivery at the various tiers of the infrastructure continuum (access, edge, regional data center, cloud). We propose a distributed dynamic policy for ML model allocation in an IDNby which each node dynamically updates its local set of inference models based on requests observed during the recent past plus limited information exchange with its neighboring nodes. Our policy offers strong performance guarantees in an adversarial setting and shows improvements over greedy heuristics with similar complexity in realistic scenarios.  \nI. INTRODUCTION  \nMachine learning (ML) models are often trained to perform inference, that is to elaborate predictions based on input data. ML model training is a computationally and I/O intensive operation and its streamlining is the object of much research effort. Although inference does not involve complex iterative algorithms and is therefore generally assumed to be easy, it also presents fundamental challenges that are likely to become dominant as ML adoption increases [1] . In a future where AI systems are ubiquitously deployed and need to make timely and safe decisions in unpredictable environments, inference requests will have to be served in real-time and the aggregate rate of predictions needed to support a pervasive ecosystem of sensing devices will become overwhelming.  \nToday, two deployment options for ML models are common: inferences can be served by the end devices (smartphones, IoT equipment, smart vehicles, etc.), where only simple models can run, or by a remote cloud infrastructure, where powerful “machine learning as a service” (MLaaS) solutions rely on sophisticated models and provide inferencesat extremely high throughput.  \nHowever, there exist applications for which both options may be unsuitable: local models may have inadequate ac-  \ncuracy, while the cloud may fail to meet delay constraints. As an example, popular applications such as recommendation systems, voice assistants, and ad-targeting, need to serve predictions from ML models in less than 200 ms. Future wireless services, such as connected and autonomous cars, industrial robotics, mobile gaming, augmented/virtual reality, have even stricter latency requirements, often below 10 ms and","cbCaipwKfBnqvi6B","https://ap.wps.com/l/cbCaipwKfBnqvi6B","pdf",1131347,1,29,"English","en",105,"# Introduction\n## Deployment options for ML inference\n## Inference Delivery Networks (IDNs)\n## Optimization problem for model allocation","[{\"question\":\"What problem does the paper address for ML inference deployment?\",\"answer\":\"It addresses cases where local execution cannot meet accuracy targets and cloud execution cannot satisfy strict delay constraints, motivating a better way to place and select models across infrastructure tiers.\"},{\"question\":\"What is an inference delivery network (IDN)?\",\"answer\":\"An IDN is a network of computing nodes that coordinate to serve ML inference requests, bridging device and cloud execution by integrating inference delivery across access, edge, regional data centers, and cloud.\"},{\"question\":\"How does the proposed policy allocate ML models in an IDN?\",\"answer\":\"Each node dynamically updates its local inference model set based on observed requests from the recent past and limited information exchange with neighboring nodes.\"}]","Towards Inference Delivery Networks - Distributing Machine Learning with Optimality Guarantees | PDF",1785899680,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-inference-delivery-networks-distributing-machine-learning-with-optimality-guarantees","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-inference-delivery-networks-distributing-machine-learning-with-optimality-guarantees/125530/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address for ML inference deployment?","Question",{"text":75,"@type":76},"It addresses cases where local execution cannot meet accuracy targets and cloud execution cannot satisfy strict delay constraints, motivating a better way to place and select models across infrastructure tiers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is an inference delivery network (IDN)?",{"text":80,"@type":76},"An IDN is a network of computing nodes that coordinate to serve ML inference requests, bridging device and cloud execution by integrating inference delivery across access, edge, regional data centers, and cloud.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed policy allocate ML models in an IDN?",{"text":84,"@type":76},"Each node dynamically updates its local inference model set based on observed requests from the recent past and limited information exchange with neighboring nodes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]