[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126866-en":3,"doc-seo-126866-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126866,1099523885336,"Violet","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Comparison of Autoscaling Frameworks for Containerised Machine-Learning Applications in a Local and Cloud Environment - read online free","Automated allocation of computing resources, or autoscaling, is essential for deploying machine learning (ML) inference workloads while keeping stable response times under fluctuating demand. The study evaluates deployment techniques across application-level scaling (TorchServe, RayServe) and container-level scaling (K3s) in a local environment, then compares container and machine-level scaling in AWS using ECS and EKS. Performance is assessed via mean and standard deviation of inference time for a multi-client scenario and by upscaling response times, leading to a recommended strategy for local and cloud deployments.","Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment  \n1st Christian Schrder Manufacturing Technologies Vitesco Technologies GmbH Limbach-Oberfrohna, Germany [cschroeder.research@gmail.com](cschroeder.research@gmail.com)  \n2nd Ren Bhm  \nManufacturing Technologies Vitesco Technologies GmbH Limbach-Oberfrohna, Germany [rene.boehm@vitesco.com](rene.boehm@vitesco.com)  \n3rd Alexander Lampe Dept. Engineering University of Applied Sciences Mittweida, Germany [lampe@hs-mittweida.de](lampe@hs-mittweida.de)  \narXiv :2311 . 18659v2 [ cs .DC] 25 Feb 2024  \nAbstract—When deploying machine learning (ML) applications, the automated allocation of computing resources - commonly referred to as autoscaling - is crucial for maintaining a consistent inference time under fluctuating workloads. The objective is to maximize the Quality of Service metrics, emphasizing performance and availability, while minimizing resource costs. In this paper, we compare scalable deployment techniques across three levels of scaling: at the application level (TorchServe, RayServe) and the container level (K3s) in a local environment (production server), as well as at the container and machine levels in a cloud environment (Amazon Web Services Elastic Container Service and Elastic Kubernetes Service). The comparison is conducted through the study of mean and standard deviation of inference time in a multi-client scenario, along with upscaling response times. Based on this analysis, we propose a deployment strategy for both local and cloud-based environments.  \nIndex Terms—Autoscaling, Cloud computing, Kubernetes, Amazon Web Services, Scalability, Container, Machine Learning Inference  \nI. INTRODUCTION  \nML models can be utilized for anomaly detection in automated optical quality controls or time series data analysis for manufacturing processes such as screwing, welding, or mounting. The deployment of these models must adhere to Quality of Service requirements, encompassing performance metrics like inference times and availability metrics related to upscaling response times. Additional challenges [1] arise during the transition of software from the development environment to the production environment. The IT department’s support for software solutions in a production setting necessitates standardized, production-tested frameworks. With 1,416 logos featured on the 2023 Machine Learning, Artificial Intelligence, and Data Landscape [2], establishing a standard proves to bea significant challenge.  \nAutoscaling has become a highly researched field, with a predominant focus on the development of autoscaling rules. Common techniques involve the use of machine learning algorithms such as reinforcement learning [3] [4] [5] [6], decision trees [7] [8] or long short term models [9] . Other methods are based on multi-level metric monitoring [10], fuzzy logic [11], second order autoregressive moving average  \n(ARMA) [12] and simple moving average (SMA) [13] . Additionally, combinations of local predictors, including support vector machines and ARMA [14] or long short term memory neural networks and autoregressive integrated moving average [15] are employed. While most autoscaling techniques focus on horizontal scaling, vertical scaling [16] and combinations of both [17] [18] are also of great interest. Reference [19] compares the autoscaling services of three cloud provider (Amazon Web Services, Microsoft Azure, Google Cloud Platform) on virtual machine (VM) and container level. To our best knowledge, there has been no research conducted to compare container and application-level autoscaling in an ML-specific task in both local and cloud environments.  \nThis paper focuses on the use of well-known server and container orchestration frameworks. These frameworks include ML-specific webservers equipped with autoscaling capabilities, such as TorchServe and RayServe. Additionally, scalable container frameworks like Kubern","cbCaibPHwcBVpToO","https://ap.wps.com/l/cbCaibPHwcBVpToO","pdf",369364,1,6,"English","en",105,"# Introduction\n# Deployment Methods\n## Multi-model application scenario\n## Scaling levels and frameworks\n# Method Comparison\n## Multi-client evaluation\n# RayServe Performance Analysis\n## Resource usage and suitability\n# Summary and Next Steps","[{\"question\":\"What problem does the paper address when deploying ML applications?\",\"answer\":\"It addresses how to use autoscaling to keep inference time consistent under fluctuating workloads, while balancing Quality of Service with resource cost.\"},{\"question\":\"Which frameworks are compared across local and cloud environments?\",\"answer\":\"The local comparison covers TorchServe and RayServe at the application level and K3s at the container level, while the cloud comparison uses AWS ECS and EKS across container and machine levels.\"},{\"question\":\"How is performance evaluated in the comparison?\",\"answer\":\"Inference time stability is measured using the mean and standard deviation across a multi-client scenario, and scaling behavior is assessed using upscaling response times.\"}]","Comparison of Autoscaling Frameworks for Containerised Machine-Learning Applications in a Local and Cloud Environment - read online free | PDF",1785935306,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"comparison-of-autoscaling-frameworks-for-containerised-machine-learning-applications-in-a-local-and-cloud-environment-read-online-free","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/comparison-of-autoscaling-frameworks-for-containerised-machine-learning-applications-in-a-local-and-cloud-environment-read-online-free/126866/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address when deploying ML applications?","Question",{"text":75,"@type":76},"It addresses how to use autoscaling to keep inference time consistent under fluctuating workloads, while balancing Quality of Service with resource cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which frameworks are compared across local and cloud environments?",{"text":80,"@type":76},"The local comparison covers TorchServe and RayServe at the application level and K3s at the container level, while the cloud comparison uses AWS ECS and EKS across container and machine levels.",{"name":82,"@type":73,"acceptedAnswer":83},"How is performance evaluated in the comparison?",{"text":84,"@type":76},"Inference time stability is measured using the mean and standard deviation across a multi-client scenario, and scaling behavior is assessed using upscaling response times.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]