[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127508-en":3,"doc-seo-127508-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127508,13056712833777,"Logic","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","GPU-enabled Function-as-a-Service for Machine Learning Inference - paper - 2023","Function-as-a-Service (FaaS) improves scalability and usability for cloud workloads, particularly Machine-Learning (ML) inference, but current FaaS offerings lack direct GPU support. GPUs are essential for high-performance inference, yet enabling GPUs in event-triggered, short-lived functions introduces overhead from transferring model parameters and inputs/outputs. The work proposes a GPU-enabled FaaS design extending OpenFaaS with GPU-aware scheduling and GPU-memory model caching, plus locality-aware co-designed scheduling and cache management to reduce latency and maximize reuse.","GPU-enabled Function-as-a-Service for Machine  \nLearning Inference  \nMing Zhao  \nArizona State University [mingzhao@asu.edu](mingzhao@asu.edu)  \nKritshekhar Jha  \nArizona State University [kjha9@asu.edu](kjha9@asu.edu)  \nSungho Hong  \nArizona State University[shong59@asu.edu](shong59@asu.edu)  \narXiv :2303 .05601v1 [ cs .DC] 9 Mar 2023  \nAbstract—Function-as-a-Service (FaaS) is emerging as an important cloud computing service model as it can improve thescalability and usability of a wide range of applications, especially Machine-Learning (ML) inference tasks that require scalable resources and complex software conﬁgurations. These inference tasks heavily rely on GPUs to achieve high performance; however, support for GPUs is currently lacking in the existing FaaS solutions. The unique event-triggered and short-lived nature of functions poses new challenges to enabling GPUs on FaaS, which must consider the overhead of transferring data (e.g., ML model parameters and inputs/outputs) between GPU and host memory. This paper proposes a novel GPU-enabled FaaS solution that enables ML inference functions to efﬁciently utilize GPUs to accelerate their computations. First, it extends existing FaaS frameworks such as OpenFaaS to support the scheduling and execution of functions across GPUs in a FaaS cluster. Second, it provides caching of ML models in GPU memory to improve the performance of model inference functions and global management of GPU memories to improve cache utilization. Third, it offers co-designed GPU function scheduling and cache management to optimize the performance of ML inference functions. Speciﬁcally, the paper proposes locality-aware scheduling, which maximizes the utilization of both GPU memory for cache hits and GPU cores for parallel processing. A thorough evaluation based on real-world traces and ML models shows that the proposed GPU-enabled FaaS works well for ML inference tasks, and the proposed locality-aware scheduler achieves a speedup of 48x compared to the default, load balancing only schedulers.  \nIndex Terms—Function-as-a-Service, GPU scheduling, Caching, Machine learning inference  \nI. INTRODUCTION  \nFunction-as-a-Service (FaaS) has emerged as a new cloud computing service model which allows users to conveniently deploy and rapidly scale their computing tasks cost effectively. However, running machine learning (ML) inference with FaaS functions is limited as the current cloud providers do not support or directly provide FaaS functions to access GPU resources which are critical to accelerate the compute-intensive inference tasks. For example, AWS Lambda [3] does not provide GPUs to FaaS; Azure functions [9], FaaS from Microsoft can indirectly access GPUs via GPU-enabled Kubernetes containers, but they cannot share the GPUs. Therefore, there isan urgent need to enable GPUs on FaaS platforms to allow a wide variety of tasks, including ML inference, to beneﬁt from this service model.  \nThis work is partly supported by National Science Foundation awards CNS- 1955593 and OAC-2126291 .  \nEnabling GPUs on FaaS platforms is imperative to ML inference tasks as they are compute intensive and require low latency to meet the Service Level Agreement (SLA) . ML inference applications in production have stringent latency requirements; for example, providing auto-suggestions in the search bar requires returning the inference results in real-time while users browse for keywords [9] . Using GPUs to run ML inference can signiﬁcantly reduce the latency when input data can be grouped into a large batch and models are designed for parallel computation. Taking a batch of input together allows an ML model to take advantage of GPU's parallelism to process them in parallel. ML models such as Transformers [27] translate the sequential computation of recurrent neural networks (RNN) into independent calculations to beneﬁt from GPU parallelization.  \nManaging the GPU resources in a FaaS platform is challenging as sharing GPUs diffe","cbCaitOOTA8vak6f","https://ap.wps.com/l/cbCaitOOTA8vak6f","pdf",437808,3,1,11,"English","en",105,"# Introduction\n## Motivation: GPU support gap in FaaS\n## Why latency and batching matter for ML inference\n## Challenges: locality vs load balancing for GPUs\n# Proposed approach (high level)","[{\"question\":\"Why is GPU support important for running ML inference with FaaS?\",\"answer\":\"ML inference is compute intensive and latency-sensitive, and GPUs significantly reduce latency when inputs can be batched and parallelized. Existing FaaS platforms often do not provide direct access to GPUs, limiting inference performance.\"},{\"question\":\"What challenges arise when enabling GPUs on FaaS?\",\"answer\":\"Event-triggered, short-lived functions require careful handling of GPU resource management and overhead from transferring model parameters and data between GPU and host memory. Sharing GPUs in a FaaS setting is also more complex than sharing CPUs or memory.\"},{\"question\":\"How does the proposed solution improve ML inference performance on GPU-enabled FaaS?\",\"answer\":\"It extends an FaaS framework to schedule and execute functions across GPUs, caches ML models in GPU memory, and coordinates locality-aware GPU function scheduling with cache management to increase cache hits and parallel utilization, achieving up to 48x speedup versus default schedulers.\"}]","GPU-enabled Function-as-a-Service for Machine Learning Inference - paper - 2023 | PDF",1785939536,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"gpu-enabled-function-as-a-service-for-machine-learning-inference-paper-2023","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/gpu-enabled-function-as-a-service-for-machine-learning-inference-paper-2023/127508/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is GPU support important for running ML inference with FaaS?","Question",{"text":76,"@type":77},"ML inference is compute intensive and latency-sensitive, and GPUs significantly reduce latency when inputs can be batched and parallelized. Existing FaaS platforms often do not provide direct access to GPUs, limiting inference performance.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What challenges arise when enabling GPUs on FaaS?",{"text":81,"@type":77},"Event-triggered, short-lived functions require careful handling of GPU resource management and overhead from transferring model parameters and data between GPU and host memory. Sharing GPUs in a FaaS setting is also more complex than sharing CPUs or memory.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed solution improve ML inference performance on GPU-enabled FaaS?",{"text":85,"@type":77},"It extends an FaaS framework to schedule and execute functions across GPUs, caches ML models in GPU memory, and coordinates locality-aware GPU function scheduling with cache management to increase cache hits and parallel utilization, achieving up to 48x speedup versus default schedulers.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]