[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120377-en":3,"doc-seo-120377-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120377,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Exploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks - Master’s thesis","Machine learning increasingly supports applications across industries, yet some deployments demand low latency that restricts available hardware choices. This master’s thesis evaluates CPUs as an alternative to GPUs for real-time, latency-critical computer vision by applying model compression. CPU performance is improved using pruning and quantization via SparseML, then executed with the DeepSparse runtime, while the GPU baseline uses TensorRT with quantization. Results indicate CPU methods can outperform GPUs in certain scenarios, enabling latency-sensitive vision beyond GPU-limited settings.","Exploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks  \nAn evaluation of CPUs as an alternative hardware for real-time computer vision applications by using model compression  \nMaster’s thesis in Complex Adaptive Systems  \nMAX SEDERSTEN AMANDA SIKLUND  \nDEPARTMENT OF PHYSICS  \nCHALMERS UNIVERSITY OF TECHNOLOGY Gothenburg, Sweden 2024  \n[www.chalmers.se](www.chalmers.se)  \nMaster’s thesis 2024  \nExploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks  \nAn evaluation of CPUs as an alternative hardware for real-time computer vision applications by using model compression  \nMAX SEDERSTEN  \nAMANDA SIKLUND  \nDepartment of Physics Chalmers University of Technology Gothenburg, Sweden 2024  \nExploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks An evaluation of CPUs as an alternative hardware for real-time computer vision applications by using model compression  \nMAX SEDERSTEN AMANDA SIKLUND  \n© MAX SEDERSTEN, AMANDA SIKLUND, 2024 .  \nSupervisor: Filip Wikman, Tenfifty  \nExaminer: Mats Granath, Department of Physics  \nMaster’s Thesis 2024 Department of Physics  \nChalmers University of Technology SE-412 96 Gothenburg Telephone +46 31 772 1000  \nTypeset in LATEX  \nPrinted by Chalmers Reproservice Gothenburg, Sweden 2024  \nExploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks An evaluation of CPUs as an alternative hardware for real-time computer vision  \napplications by using model compression MAX SEDERSTEN, AMANDA SIKLUND Department of Physics  \nChalmers University of Technology  \nAbstract  \nIn recent years, machine learning has grown to become increasingly prevalent for a wide range of applications spanning multiple industries. For some of these applications, low latency can be critical, which may limit the types of hardware that can be used. Graphical Processing Units (GPUs) have long been the go-to hardware for machine learning tasks, often outperforming alternatives like Central Processing Units (CPUs), but these are not practical in all situations. We explore CPUs, leveraging modern optimization techniques like pruning and quantization, as a competitive alternative to GPUs with comparable predictive performance. This thesis provides a comparison of the two hardware types on a real-time latency-critical vision task. On the GPU side, TensorRT in combination with quantization is used to achieve state-of-the-art inference performance on the hardware. On the CPU side, the model is optimized using SparseML to introduce unstructured sparsity and quantization. This optimized model is then used by the DeepSparse runtime engine for optimized inference. Our findings show that the CPU approach can outperform the GPU hardware in certain situations. This suggests that CPU hardware could potentially be used in applications previously limited to GPUs.  \nKeywords: machine learning, neural network, model compression, pruning, quantization, optimization, CPU, GPU, Neural Magic, NVIDIA  \nAcknowledgements  \nWe would like to thank our supervisor at Tenfifty, Filip Wikman, for his guidance and support throughout this thesis. His expertise and insightful contributions have been a valuable part of shaping the direction and outcomes of our work.  \nWe would also like to thank our supervisor and examiner at Chalmers, Mats Granath, for his valuable feedback and assistance, particularly in providing insightful guidance that helped us shape the project outline.  \nMax Sedersten and Amanda Siklund, Gothenburg, June 2024  \nList of Acronyms  \nBelow is the list of acronyms that have been used throughout this thesis listed in alphabetical order:  \nAI Artificial Intelligence  \nASIC Application-Specific Integrated Circuit  \nAP Average Precision  \nCNN Convolutional Neural Network  \nCPU Central Processing Unit  \nGPU Graphical Processing Unit  \nIoU Intersection over Union  \nmAP mean Average Precision  \nOBD Optimal Brain Damage  \nOBS Optimal Brain Surgeon  \nOKS Object Keypoint Simi","cbCair2ECA33ZQxl","https://ap.wps.com/l/cbCair2ECA33ZQxl","pdf",2040175,1,61,"English","en",105,"# Abstract\n# Keywords\n# Acknowledgements\n# List of Acronyms","[{\"question\":\"Why does this thesis compare CPUs and GPUs for latency-critical machine learning tasks?\",\"answer\":\"Low latency can be essential in real-time applications, which limits usable hardware. The thesis evaluates whether CPUs can compete with GPUs under those constraints.\"},{\"question\":\"How is the CPU-optimized model created and executed?\",\"answer\":\"The model is optimized with SparseML to introduce unstructured sparsity and quantization, and then run using the DeepSparse runtime engine for optimized inference.\"},{\"question\":\"What hardware and software stack is used for the GPU baseline?\",\"answer\":\"The GPU side uses TensorRT combined with quantization to reach state-of-the-art inference performance on the target hardware.\"}]","Exploring Optimized CPU-Inference for Latency-Critical Machine Learning Tasks - Master’s thesis | PDF",1785729731,154,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"exploring-optimized-cpu-inference-for-latency-critical-machine-learning-tasks-masters-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exploring-optimized-cpu-inference-for-latency-critical-machine-learning-tasks-masters-thesis/120377/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does this thesis compare CPUs and GPUs for latency-critical machine learning tasks?","Question",{"text":75,"@type":76},"Low latency can be essential in real-time applications, which limits usable hardware. The thesis evaluates whether CPUs can compete with GPUs under those constraints.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the CPU-optimized model created and executed?",{"text":80,"@type":76},"The model is optimized with SparseML to introduce unstructured sparsity and quantization, and then run using the DeepSparse runtime engine for optimized inference.",{"name":82,"@type":73,"acceptedAnswer":83},"What hardware and software stack is used for the GPU baseline?",{"text":84,"@type":76},"The GPU side uses TensorRT combined with quantization to reach state-of-the-art inference performance on the target hardware.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]