[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119528-en":3,"doc-seo-119528-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119528,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Classifying Performance Bounds Using Machine Learning - Poster","Traditional performance models such as the Roofline approach rely on visual judgment to determine the correct performance bound, which becomes difficult on modern HPC CPUs featuring deep cache hierarchies and front-end out-of-order execution. This poster presents an initial, data-driven performance modelling workflow based on machine learning. Supervised and unsupervised models are trained and evaluated on a curated benchmark dataset built from well-understood applications, using normalized performance counters as features.","Littman, L. , & Deakin, T. (2025) . Classifying Performance Bounds Using Machine Learning. Poster session presented at Supercomputing, St Louis, Missouri, United States.  \nPeer reviewed version  \nLink to publication record on the Bristol Research Portal  \nPDF-document  \nThis is the accepted author manuscript (AAM) . The final published version (Version of Record) can be found on the publisher's website. The copyright of any third-party content, such as images, remains with the copyright holder.  \nUniversity of Bristol – Bristol Research Portal  \nGeneral rights  \nThis document is made available in accordance with publisher policies. Please cite only the published version using the reference above. Full terms of use are available: [http://www.bristol.ac.uk/red/research-policy/pure/user-guides/brp-terms/](http://www.bristol.ac.uk/red/research-policy/pure/user-guides/brp-terms/)  \nClassifying Performance Bounds Using Machine Learning  \nLewis Littman and Tom Deakin  \nUniversity of Bristol, Bristol, UK  \nAbstract  \nTraditional performance analysis tools, such as the Roofline model, require visual interpretation to determine performance bounds. For CPUs which have complex cache hierarchies and front-end out-of-order capabilities—that is the CPUs we use for high performance computing—accurately identifying the true performance bound is challenging. This work is the first steps towards a data-driven approach to performance modelling, leveraging Machine Learning techniques. We build and evaluate a number of supervised and unsupervised models using a new curated data set of performance counters collected from well-understood (i.e., easily labeled) benchmark applications. We further analyse the data set and highlight potential “performance fingerprints” obtainable using this methodology.  \nPerformance Limiting Factors  \nThe Roofline model (Williams, Waterman and Patterson) categorises the performance limiting factor of an application/kernel by considering how its performance (in FLOPS) compares to theoretical peak on that processor. In particular, based on the Arithmetic Intensity—the ratio of floating-point operations to data movement—that kernel will be classed as “compute bound” or “main memory bandwidth bound” based on which side of the ridge point it falls. In order for the categorisation to be considered correct the performance should be close to the roofline bound, often cited as 50% .  \nPerformance FLOP/s log-scale  \nMemory bandwidth bound  \nPeak FLOP/s  \n\"Ridge point\"  \nFloating-point bound  \nlog-scale  \nOperational intensity FLOPs/byte  \nAccess the data set  \nData Set generation  \nWe collected a data set of performance counters by executing each of the following applications/problem sizes. Each application was run 3 times to collect all performance counters, and together they form a single record in the data set. 100 records for each application/problem size were collected, resulting in 1,200 records in the new data set. Performance bounds were verified by Roofline analysis even though they are well known.  \nBenchmark  \nSGEMMDGEMM miniBUDE STREAM LBM D2Q9 MiniFE  \n3D Heat stencil  \nProblem Sizes  \n15kx-by-15k and 10k-by-10k  \n15k-by-15k and 10k-by-10k 128 PPWI, WGS 1  \n10M elements  \n256-by-256 and 1024-by-1024 50-by-50-by-50 120-by-120-by-120  \nPerformance Bound  \nCompute  \nCompute  \nCompute  \nMain Memory Bandwidth LLC Memory Bandwidth Main Memory Bandwidth LLC memory bandwidth  \nLU Solver (MKL) 8192-by-8192 and 16384-by-16384 LLC memory bandwidth  \nPerformance counters were collected using Linux perf and normalised. Each record in the data set therefore contains:  \n• GFLOPs  \n• FLOPc  \n• IPC  \n• retiring %  \n• bad speculation %  \n• frontend bound %  \n• backend bound %  \n• vectorised SP instruction ratio  \n• vectorised DP instruction  \nratio  \n• cache miss ratio  \n• L1 cache miss ratio  \n• L2 cache miss ratio and L3 cache miss ratio  \nBaseline Models  \nThree naive models were tested to create a baseline for the machine learned ","cbCaihlMJ4P6mFfc","https://ap.wps.com/l/cbCaihlMJ4P6mFfc","pdf",386283,1,2,"English","en",105,"# Abstract\n# Performance Limiting Factors\n## Roofline model and ridge point\n# Data Set Generation\n## Benchmarks and problem sizes\n## Collected performance counters\n# Baseline Models\n# Machine Learning Models\n## Models and evaluation\n## t-SNE analysis","[{\"question\":\"Why is identifying the true performance bound difficult on HPC CPUs?\",\"answer\":\"Modern HPC CPUs have complex cache hierarchies and front-end out-of-order capabilities, which makes performance bounds harder to determine accurately using traditional visual methods.\"},{\"question\":\"How does the Roofline model classify performance bounds?\",\"answer\":\"It compares a kernel’s FLOP/s to the processor’s theoretical peak using arithmetic intensity, labeling it as compute bound or main-memory bandwidth bound depending on which side of the ridge point it falls.\"},{\"question\":\"What data and features are used to train the machine learning models?\",\"answer\":\"A curated dataset is generated by running benchmark applications multiple times and collecting normalized Linux perf performance counters, including metrics such as GFLOPs/FLOPc, IPC, speculation and bound percentages, and cache-miss ratios.\"}]","Classifying Performance Bounds Using Machine Learning - Poster | PDF",1785724784,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"classifying-performance-bounds-using-machine-learning-poster","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":21},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/classifying-performance-bounds-using-machine-learning-poster/119528/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is identifying the true performance bound difficult on HPC CPUs?","Question",{"text":74,"@type":75},"Modern HPC CPUs have complex cache hierarchies and front-end out-of-order capabilities, which makes performance bounds harder to determine accurately using traditional visual methods.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the Roofline model classify performance bounds?",{"text":79,"@type":75},"It compares a kernel’s FLOP/s to the processor’s theoretical peak using arithmetic intensity, labeling it as compute bound or main-memory bandwidth bound depending on which side of the ridge point it falls.",{"name":81,"@type":72,"acceptedAnswer":82},"What data and features are used to train the machine learning models?",{"text":83,"@type":75},"A curated dataset is generated by running benchmark applications multiple times and collecting normalized Linux perf performance counters, including metrics such as GFLOPs/FLOPc, IPC, speculation and bound percentages, and cache-miss ratios.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":29,"slug":136},19,"General","general"]