[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117225-en":3,"doc-seo-117225-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117225,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Benchmarking Machine Learning Applications on Heterogeneous Architecture using Reframe - Preliminary Results and Kubernetes Backend Integration","Machine learning workloads are increasingly deployed on HPC systems, making regular, machine-learning-specific benchmarks essential for monitoring performance and detecting issues. This work extends the ReFrame framework to add a Kubernetes scheduler backend, enabling repeatable benchmarking across diverse platforms. The study targets EPCC’s ML accelerator environment within the Edinburgh International Data Facility, covering Nvidia GPUs and novel accelerators such as Graphcore Bow Pod64 and Cerebras CS-2, and discusses preliminary results and integration challenges.","arXiv :2404 . 10536v2 [ cs .DC] 25 Apr 2024  \nBenchmarking Machine Learning Applications on Heterogeneous Architecture using Reframe  \nChristopher Rae  \nJoseph K. L. Lee  \nJames Richings  \nMichèle Weiland  \n[crae@ed.ac.uk](crae@ed.ac.uk)  \n[j.lee@epcc.ed.ac.uk](j.lee@epcc.ed.ac.uk)  \n[j.richings@epcc.ed.ac.uk](j.richings@epcc.ed.ac.uk)  \n[m.weiland@epcc.ed.ac.uk](m.weiland@epcc.ed.ac.uk)  \nEPCC, The University of Edinburgh  \nUK  \nABSTRACT  \nWith the rapid increase in machine learning workloads performed on HPC systems, it is beneﬁcial to regularly perform machine learning speciﬁc benchmarks to monitor performance and identify issues. Furthermore, as part of the Edinburgh International Data Facility, EPCC currently hosts a wide range of machine learning accelerators including Nvidia GPUs, the Graphcore Bow Pod64 and Cerebras CS-2, which are managed via Kubernetes and Slurm. We extended theReframe framework to support the Kubernetes scheduler backend, and utilise Reframe to perform machine learning benchmarks, and we discuss the preliminary results collected and challenges involved in integrating Reframe across multiple platforms and architectures.  \nFor example, at EPCC, besides managing HPC systems, ARCHER2 (the UK national HPC service) and Cirrus (an EPSRC tier-2 system), a new collection of data science-focused services that come under the umbrella of the Edinburgh International Data Facility (EIDF) [8] has been created to support research and data-driven innovation for academic and commercial partners. As part of the EIDF service, we oﬀer access to a GPU cluster, large memory system (HPE Superdome Flex), as well as dedicated ML accelerators including the Cerebras CS-2 [4]and Graphcore Bow Pod64 [17]. To use the EIDF GPU service and the Graphcore system, Kubernetes is used to manage workload, instead of a HPC scheduler such as Slurm [29]. This is to allow data scientists to run their applications in a reproducible and portable container environment, with maximum ease and minimal setup required.  \nCCS CONCEPTS With the increasingly diverse range of services being oﬀered, it  \nis critical to regularly perform regression testing and benchmark-  \n• Software and its engineering → Software performance.  \ning to maintain quality of service and detect issues. This allows  \nKEYWORDS us to monitor hardware performance and software compatibility,  \ntrack performance around system changes or updates, to ensure  \nHPC, Reframe, Kubernetes, Benchmarking, Data Science, Machine service quality. The aim of this work is to create a framework for learning, MLPerf, GPU, Graphcore, Cerebras repeatable testing of multiple hardware architectures at EPCC.  \nACM Reference Format: The main contributions of this paper are:  \nChristopher Rae, Joseph K. L. Lee, James Richings, and Michèle Weiland.  \n􀀏 Integrate Kubernetes as a backend for the Reframe [12] test-  \n2024. Benchmarking Machine Learning Applications on Heterogeneous Ar  \nchitecture using Reframe. In The 33rd International Symposium on High- ing framework, and explain the steps required to set up test  \nPerformance Parallel and Distributed Computing (HPDC ’24), June 3–7, 2024, cases  \nPisa, Italy. ACM, New York, NY, USA,7pages. [https://doi.org/10.1145/3625549.3658896](https://doi.org/10.1145/3625549.3658896) 􀀏 Demonstrate and compare ML benchmarks (ResNet-50, Deep-  \n1 INTRODUCTION  \nTo support the increasing trend of Machine Learning (ML) applications being used inscientiﬁc research and taking advantage ofHPC systems, national HPC providers are adapting in two ways: 1) by providing compute services more suitable for data scientists which are more akin to cloud platforms, and 2) adoption of dedicated ML accelerators besides traditional HPC hardware (CPUs and GPUs) .  \nThis is the author’s version of the work. The deﬁnitive version was published in The 33rd International Symposium on High-Performance Parallel and Distributed Computing (HPDC ’24), June3–7, 2024, Pisa, Italy, [https://doi.org/10","cbCaipf671ofugXl","https://ap.wps.com/l/cbCaipf671ofugXl","pdf",173113,1,7,"English","en",105,"# Introduction\n# Background\n## Reframe\n# Benchmarking Setup and Kubernetes Backend\n# Preliminary Results and Challenges","[{\"question\":\"Why are machine learning benchmarks important for HPC systems?\",\"answer\":\"Regular machine learning-specific benchmarks help monitor performance, track compatibility issues, and detect regressions after system changes or updates.\"},{\"question\":\"What extension was added to ReFrame in this work?\",\"answer\":\"The authors extended ReFrame to support a Kubernetes scheduler backend, allowing benchmarks to run in a containerized, reproducible workflow.\"},{\"question\":\"Which platforms and ML accelerators are targeted in the benchmarking environment?\",\"answer\":\"The work focuses on EPCC’s accelerator platforms within EIDF, including Nvidia GPUs and dedicated accelerators such as Graphcore Bow Pod64 and Cerebras CS-2, managed with Kubernetes rather than an HPC scheduler like Slurm.\"}]","Benchmarking Machine Learning Applications on Heterogeneous Architecture using Reframe - Preliminary Results and Kubernetes Backend Integration | PDF",1785674489,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"benchmarking-machine-learning-applications-on-heterogeneous-architecture-using-reframe-preliminary-results-and-kubernetes-backend-integration","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/benchmarking-machine-learning-applications-on-heterogeneous-architecture-using-reframe-preliminary-results-and-kubernetes-backend-integration/117225/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are machine learning benchmarks important for HPC systems?","Question",{"text":75,"@type":76},"Regular machine learning-specific benchmarks help monitor performance, track compatibility issues, and detect regressions after system changes or updates.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What extension was added to ReFrame in this work?",{"text":80,"@type":76},"The authors extended ReFrame to support a Kubernetes scheduler backend, allowing benchmarks to run in a containerized, reproducible workflow.",{"name":82,"@type":73,"acceptedAnswer":83},"Which platforms and ML accelerators are targeted in the benchmarking environment?",{"text":84,"@type":76},"The work focuses on EPCC’s accelerator platforms within EIDF, including Nvidia GPUs and dedicated accelerators such as Graphcore Bow Pod64 and Cerebras CS-2, managed with Kubernetes rather than an HPC scheduler like Slurm.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]