[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118046-en":3,"doc-seo-118046-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118046,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","A Survey of Serverless Machine Learning Model Inference","Generative AI, computer vision, and NLP adoption is driving machine learning models into production, where real-time inference must satisfy SLOs for reliability, low downtime, and cost control. Many large models rely on GPU resources to achieve acceptable latency, yet serverless hosting often lacks GPU support or incurs continuous 24/7 billing with GPU access. This survey summarizes emerging challenges and optimization opportunities for large-scale deep learning serving systems, providing a taxonomy and highlighting trends that guide future inference framework research.","A Survey of Serverless Machine Learning Model Inference  \nKamil Kojs  \nIT University of Copenhagen, Denmark  \n[kako@itu.dk](kako@itu.dk)  \narXiv :2311 . 13587v1 [ cs .DC] 22 Nov 2023  \nAbstract  \nRecent developments in Generative AI, Computer Vi sion, and Natural Language Processing have led to an increased integration of AI models into various products. This widespread adoption of AI requires signi􀀂cant efforts in deploying these models in production environments. When hosting machine learning models for real-time predictions, it is important to meet de􀀂ned Service Level Objectives (SLOs), ensuring reliability, minimal downtime, and opti mizing operational costs of the underlying infrastructure. Large machine learning models often demand GPU resources for ef􀀂cient inference to meet SLOs. In the context of these trends, there is growing interest in hosting AI models in a serverless architecture while still providing GPU access for inference tasks. This survey aims to summarize and categorize the emerging challenges and optimization opportunities for large-scale deep learning serving systems. By providing a novel taxonomy and summarizing recent trends, we hope that this survey could shed light on new optimization perspectives and motivate novel works in large-scale deep learning serving systems.  \n1. Introduction  \nModel serving systems are designed to handle user inference requests in real time. This interactive nature differentiates them from model training systems, which are primarily focused on maximizing data processing throughput. The design of model serving systems is driven by several key goals. Firstly, high performance is crucial, as the system needs to process requests swiftly, even under variable and high-demand workloads. Secondly, cost-effectiveness is a major consideration, as the system should be able to handle a large volume of requests without incurring excessive costs. Thirdly, ease of management is important, allowing data scientists to deploy machine learning models ef􀀂ciently without being slowed down by intricate details of resource management.  \nCurrently, cloud service providers offer a range of model serving options, including both managed machine learning  \nservices and self-managed server solutions. However, these offerings often do not fully satisfy the aforementioned objectives. Available options include serverless deployments that lack GPU support, where billing is based on the duration of system usage [14] [5] [7] . This model, while seemingly ef􀀂cient for non-machine learning applications, can lead to increased latency in response times, potentially causing failures in meeting established SLOs due to its lack of GPU support. On the other hand, there are deployments that provide GPU access for machine learning models, but these typically involve continuous billing throughout the entire lifecycle of the infrastructure [15] [4] [6] . This means that costs are incurred 24/7 for online serving systems, significantly increasing the operational expenses for businesses that offer AI solutions.  \nRecent developments and trends show that the model parameter sizes are growing substantially with every new generation of developed models. Notable examples include OpenAI’s GPT-3.5 and its successor GPT-4 with parameter sizes equivalent to 175 billion and 1.76 trillion respectively [29] . The big parameter size and complexity of these models make them unsuitable for environments that do not have access to GPU resources. Generally, any solution that incorporates large deep learning models and necessitates frequent machine learning inference is contingent on the availability of GPU resources. This requirement underscores the critical role of GPUs in the ef􀀂cient and effective deployment of advanced deep-learning models.  \nAnother critical aspect of deploying machine learning models for inference tasks is the suboptimal utilization of the underlying GPU infrastructure. While GPUs are typically used ef􀀂ciently duri","cbCaiqxEp7Q0SFYv","https://ap.wps.com/l/cbCaiqxEp7Q0SFYv","pdf",186676,1,13,"English","en",105,"# Introduction\n# Methods","[{\"question\":\"Why is model serving different from model training?\",\"answer\":\"Model serving processes user inference requests in real time and must handle variable, high-demand workloads. Training mainly focuses on maximizing data processing throughput rather than interactive latency.\"},{\"question\":\"What are the main limitations of current serverless model hosting for inference?\",\"answer\":\"Some serverless options lack GPU support, which can increase response latency and risk violating SLOs. Others provide GPU access but typically charge continuously throughout the infrastructure lifecycle, raising operational costs.\"},{\"question\":\"Why do large machine learning models need GPUs for inference?\",\"answer\":\"As model parameter sizes and complexity grow, inference in production generally becomes impractical without GPU resources. GPU availability is therefore critical to efficient and effective deep-learning deployment.\"}]","A Survey of Serverless Machine Learning Model Inference | PDF",1785681001,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-survey-of-serverless-machine-learning-model-inference","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-survey-of-serverless-machine-learning-model-inference/118046/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is model serving different from model training?","Question",{"text":76,"@type":77},"Model serving processes user inference requests in real time and must handle variable, high-demand workloads. Training mainly focuses on maximizing data processing throughput rather than interactive latency.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the main limitations of current serverless model hosting for inference?",{"text":81,"@type":77},"Some serverless options lack GPU support, which can increase response latency and risk violating SLOs. Others provide GPU access but typically charge continuously throughout the infrastructure lifecycle, raising operational costs.",{"name":83,"@type":74,"acceptedAnswer":84},"Why do large machine learning models need GPUs for inference?",{"text":85,"@type":77},"As model parameter sizes and complexity grow, inference in production generally becomes impractical without GPU resources. GPU availability is therefore critical to efficient and effective deep-learning deployment.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]