[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124036-en":3,"doc-seo-124036-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124036,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",6,"Technology","Optimizing Performance on Trinity - Utilizing Machine Learning, Proxy Applications and Scheduling Priorities","Supercomputers like Trinity scale to thousands of nodes, while real workloads progress at the rate of the slowest components. The paper addresses the ongoing need to detect underperforming nodes, assess their impact, and reduce downtime during performance-critical runs. It presents fast proxy tests based on MPI and OpenMP to replace long runtimes, then applies machine learning and clustering to isolate outliers. The resulting ordered node lists support scheduling policies that minimize slow-node effects and improve overall cluster efficiency.","arXiv :2404 . 10617v1 [ cs .DC] 16 Mar 2024  \nOptimizing Performance on Trinity Utilizing Machine Learning, Proxy Applications and  \nScheduling Priorities  \nPhil Romero  \nHigh Performance Computing Division  \nLos Alamos National Laboratory  \nEmail: [prr@lanl.gov](prr@lanl.gov)  \nAbstract—Abstract—The sheer number of nodes continues to increase in today’s supercomputers, the first half of Trinity alone contains more than 9400 compute nodes. Since the speed of today’s clusters are limited by the slowest nodes, it more important than ever to identify slow nodes, improve their performance if it can be done, and assure minimal usage of slower nodes during performance critical runs. This is an ongoing maintenance task that occurs on a regular basis and, therefore, it is important to minimize the impact upon its users by assessing and addressing slow performing nodes and mitigating their consequences while minimizing down time. These issues can be solved, in large part, through a systematic application of fast running hardware assessment tests, the application of Machine Learning, and making use of performance data to increase efficiency of large clusters. Proxy applications utilizing both MPI and OpenMP were developed to produce data as a substitute for long runtime applications to evaluate node performance. Machine learning is applied to identify underperforming nodes, and policies are being discussed to both minimize the impact of underperforming nodes and increase the efficiency of the system. In this paper, I will describe the process used to produce quickly performing proxy tests, consider various methods to isolate the outliers, and produce ordered lists for use in scheduling to accomplish this task.  \nKeywords–Machine Learning, High Performance Computing, Benchmarking, Artificial Intelligence.  \nI. INTRODUCTION  \nToday’s supercomputers consist of ever larger and growing numbers of nodes, since computational tasks running on these computers are limited to progressing at rates that are limited by the slowest performing components, it is more important than ever to identify bottlenecks and mitigate their damaging effects in a time efficient manner. There are several factors that contribute to producing the slow progression of computational tasks including:  \n1) Slow performing CPU’s  \n2) Slow performing Random Access Memory  \n3) Slow performing and/or less than optimally buffered input/output systems  \n4) Slow performing node interconnects.  \nFactory tests allowed Los Alamos extensive time to test many different applications on a two thousand node subset of Trinity. Four different tests were conducted targeting each performance category of cpu speed, memory speed, and interconnect bandwidth. The results were then mined for clusters utilizing the kMeans clustering algorithm[1] . The output of the  \nFigure 1 . This figure shows a map plot that attempts to preserve distances between the features of all items under consideration. Note the cluster labeled”3”, it shows a markedly different shape from the clusters labeled ”1” with large outliers. The x and y axes represent distances in arbitrary space and are orthogonal.  \nclustering process proved interesting in that cpu speed tests produced the largest variations in performance. This can be seen in Figure 1, the clusters labeled 2 and 3 show widespread outliers that produce irregular shape clusters as compared to the cluster labeled 1 . This plot attempts to preserve distances in two dimensions that enable a better assessment of the variability in items than would be found through a Principal Components Analysis[4] that is sensitive to a given projection of data from many dimensions into two dimensions. Similar plots produced for memory speed and interconnect bandwidth do not show widespread outliers, they produce relatively compact and regular shapes.This performance imbalance indicates that extra effort should be expended in reducing the differences in performance amongst cpu relat","cbCaijCbheoAYaRl","https://ap.wps.com/l/cbCaijCbheoAYaRl","pdf",1552870,1,7,"English","en",105,"# Introduction\n## Node bottlenecks and performance categories\n## Limitations of full HPL measurements\n# Methodology: Proxy applications\n## Outline of proxy test approach\n# Results and mapping to HPL\n## High Performance Linpack comparison\n# Mitigation and scheduling strategies\n# Conclusion","[{\"question\":\"Why is identifying slow nodes critical on Trinity-like supercomputers?\",\"answer\":\"Cluster-wide task progress is limited by the slowest nodes, so detecting bottlenecks early helps avoid damaging performance drops and reduces user-impacting downtime.\"},{\"question\":\"What problem do proxy applications solve compared with standard HPL runs?\",\"answer\":\"HPL can take more than four hours per tuned assessment and requires multiple measurements for confidence, so fast proxy tests provide quicker performance information.\"},{\"question\":\"How does machine learning contribute to improving scheduling decisions?\",\"answer\":\"Machine learning is used to identify underperforming nodes, isolate outliers, and generate ordered node lists that scheduling policies can exploit to reduce the impact of slow nodes.\"}]","Optimizing Performance on Trinity - Utilizing Machine Learning, Proxy Applications and Scheduling Priorities | PDF",1785819984,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"optimizing-performance-on-trinity-utilizing-machine-learning-proxy-applications-and-scheduling-priorities","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/optimizing-performance-on-trinity-utilizing-machine-learning-proxy-applications-and-scheduling-priorities/124036/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is identifying slow nodes critical on Trinity-like supercomputers?","Question",{"text":75,"@type":76},"Cluster-wide task progress is limited by the slowest nodes, so detecting bottlenecks early helps avoid damaging performance drops and reduces user-impacting downtime.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem do proxy applications solve compared with standard HPL runs?",{"text":80,"@type":76},"HPL can take more than four hours per tuned assessment and requires multiple measurements for confidence, so fast proxy tests provide quicker performance information.",{"name":82,"@type":73,"acceptedAnswer":83},"How does machine learning contribute to improving scheduling decisions?",{"text":84,"@type":76},"Machine learning is used to identify underperforming nodes, isolate outliers, and generate ordered node lists that scheduling policies can exploit to reduce the impact of slow nodes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,117,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":114,"show_sort_weight":115,"slug":116},"Healthcare",40,"healthcare",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":119,"show_sort_weight":120,"slug":121},8,"Research & Report",30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]