[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124441-en":3,"doc-seo-124441-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124441,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Evaluation of Model Serving Frameworks for Machine Learning","This thesis evaluates machine learning model-serving frameworks with emphasis on production-relevant performance, deployment effort, and multi-model support. The study compares TensorFlow Serving, Triton Inference Server, BentoML, TorchServe, and FastAPI, then selects TensorFlow Serving, Triton, and BentoML for practical testing based on project fit. A final architecture combines TensorFlow Serving with FastAPI, using FastAPI for preprocessing, postprocessing, and OAuth2-based authentication. Testing is conducted under CPU-bound REST API scenarios to ensure broad compatibility, while proposing GPU-enabled evaluation for further gains.","BACHELORTHESIS Divyesh Joshi  \nEvaluation of Model Serving Frameworks for Machine  Learning  \nFACULTY OF COMPUTER SCIENCE AND ENGINEERING Department of Information and Electrical Engineering  \nFakultät Technik und Informatik  \nDepartment Informations-und Elektrotechnik  \nHAMBURG UNIVERSITY OF APPLIED SCIENCES  \nHochschule für Angewandte Wissenschaften Hamburg  \nDivyesh Joshi  \nEvaluation of Model Serving Frameworks for Machine Learning  \nBachelor Thesis based on the examination and study regulations for the Bachelor of Engineering degree programme  \nBachelor of Science Information Engineering  \nat the Department of Information and Electrical Engineering of the Faculty of Engineering and Computer Science  \nof the University of Applied Sciences Hamburg Supervising examiner: Prof. Dr. Kolja Eger  \nSecond examiner: Prof. Dr. Wolfgang Renz  \nDay of delivery: 04 . October 2024  \nDivyesh Joshi  \nTitle of Thesis  \nEvaluation of Model Serving Frameworks for Machine Learning  \nKeywords  \nMachine Learning, Model Serving, TensorFlow Serving, Triton Inference Server, BentoML, FastAPI, Inference, Preprocessing, Postprocessing, REST API, Authentication, Latency, Performance, Cloud Deployment, Docker, Microsoft Azure  \nAbstract  \nThis thesis presents an evaluation of model serving frameworks for machine learning, focusing on their performance, ease of deployment, and multi-model support in real-world production environments. The frameworks evaluated include TensorFlow Serving, Triton Inference Server, BentoML, TorchServe, and FastAPI. After a comprehensive theoretical analysis, TensorFlow Serving, Triton, and BentoML were selected for practical evaluation due to their compatibility with the project’s requirements.  \nThe final system integrates TensorFlow Serving with FastAPI to create a efficient machine learning model-serving platform. In this architecture, TensorFlow Serving handles inference while FastAPI is responsible for preprocessing, postprocessing, and implementing secure authentication using OAuth2 . The system was tested under CPU-bound conditions using REST APIs to ensure broad compatibility.  \nAlthough TensorFlow Serving exhibited superior performance in terms of latency, testing on GPU-enabled hardware could potentially enhance performance across all frameworks, offering even greater improvements in inference speed and efficiency. Future work can focus on conducting more extensive testing, particularly on GPU-enabled systems.  \nDivyesh Joshi  \nThema der Arbeit  \nEvaluation von Model-Serving Frameworks für maschinelles Lernen  \nStichworte  \nMaschinelles Lernen, Model Serving, TensorFlow Serving, Triton Inference Server, BentoML, FastAPI, Inferenz, Preprocessing, Postprocessing, REST API, Authentifizierung, Latenz, Performance, Cloud-Deployment, Docker, Microsoft Azure  \nKurzzusammenfassung  \nDiese Arbeit stellt eine Evaluierung von Model serving Frameworks für maschinelles Lernen vor und konzentriert sich dabei auf deren Leistung, einfache Bereitstellung und MultiModell-Unterstützung in realen Produktionsumgebungen. Zu den evaluierten Frameworks gehören TensorFlow Serving, Triton Inference Server, BentoML, TorchServe und FastAPI. Nach einer umfassenden theoretischen Analyse wurden TensorFlow Serving, Triton und BentoML aufgrund ihrer Kompatibilität mit den Anforderungen des Projektsfür die praktische Evaluierung ausgewählt.  \nDas endgültige System integriert TensorFlow Serving mit FastAPI, um eine effiziente Plattform für maschinelles Lernen und Modellserving zu schaffen. In dieser Architektur übernimmt TensorFlow Serving die Inferenz, während FastAPI für das Preprocessing, Postprocessing und die Implementierung einer sicheren Authentifizierung mittels OAuth2 verantwortlich ist. Das System wurde unter CPU-gebundenen Bedingungen mit REST APIs getestet, um eine breite Kompatibilität zu gewährleisten.  \nObwohl TensorFlow Serving eine überlegene Leistung in Bezug auf die Latenzzeit aufwies, könnte das Testen auf GPU-fähiger Hardware ","cbCaivFTKJakp5ig","https://ap.wps.com/l/cbCaivFTKJakp5ig","pdf",1300158,1,80,"English","en",105,"# Introduction\n## Motivation\n## Goals\n## Organization of Chapters\n# Theory\n## Machine Learning Model Serving\n## Machine Learning Models\n## Data Processing\n## Security\n## Containerization and Cloud Deployment\n# Requirements\n## Functional Requirements\n## Non-Functional Requirements\n## MoSCoW Priority Classification\n# Evaluation of Frameworks and Design","[{\"question\":\"Which model-serving frameworks are evaluated in the thesis?\",\"answer\":\"The thesis evaluates TensorFlow Serving, Triton Inference Server, BentoML, TorchServe, and FastAPI, then selects TensorFlow Serving, Triton, and BentoML for practical evaluation based on project requirements.\"},{\"question\":\"What is the final proposed system architecture?\",\"answer\":\"The final system integrates TensorFlow Serving with FastAPI. TensorFlow Serving performs inference, while FastAPI handles preprocessing, postprocessing, and secure authentication using OAuth2.\"},{\"question\":\"How is the system tested and what improvements are suggested?\",\"answer\":\"The system is tested under CPU-bound conditions using REST APIs for broad compatibility. Future work recommends running more extensive tests on GPU-enabled hardware to improve inference speed and efficiency.\"}]","Evaluation of Model Serving Frameworks for Machine Learning | PDF",1785822308,202,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluation-of-model-serving-frameworks-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/evaluation-of-model-serving-frameworks-for-machine-learning/124441/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which model-serving frameworks are evaluated in the thesis?","Question",{"text":75,"@type":76},"The thesis evaluates TensorFlow Serving, Triton Inference Server, BentoML, TorchServe, and FastAPI, then selects TensorFlow Serving, Triton, and BentoML for practical evaluation based on project requirements.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the final proposed system architecture?",{"text":80,"@type":76},"The final system integrates TensorFlow Serving with FastAPI. TensorFlow Serving performs inference, while FastAPI handles preprocessing, postprocessing, and secure authentication using OAuth2.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the system tested and what improvements are suggested?",{"text":84,"@type":76},"The system is tested under CPU-bound conditions using REST APIs for broad compatibility. Future work recommends running more extensive tests on GPU-enabled hardware to improve inference speed and efficiency.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":21,"slug":99},"Literature","literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]