[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119991-en":3,"doc-seo-119991-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119991,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Machine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance","Optimizing the performance of GPU-accelerated computational fluid dynamics (CFD) simulations is essential for efficient scientific computing. This study introduces a machine learning-based autotuning method that optimizes 14 GPU-kernel scheduling parameters, including thread-block and thread configurations. Fully connected neural networks predict execution time from tuning parameters, enabling autotuning without extensive sampling. Experiments on three GPU types compare independent and combined training, demonstrating strong performance gains while using only a small fraction of the full search space.","arXiv :2306 . 14011v3 [ cs .PF] 20 Feb 2024  \nMachine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance  \nWeicheng Xue  \nKevin T. Crofton Department of Aerospace and Ocean Engineering  \nVirginia Tech  \nBlacksburg, VA, 24060  \n[weich97@vt.edu](weich97@vt.edu)  \nChristopher J. Roy  \nKevin T. Crofton Department of Aerospace and Ocean Engineering  \nVirginia Tech  \nBlacksburg, VA, 24060  \n[cjroy@vt.edu](cjroy@vt.edu)  \nAbstract  \nOptimizing the performance of computational fluid dynamics (CFD) applications accelerated by graphics processing units (GPUs) is crucial for efficient simulations. In this study, we employed a machine learning-based autotuning technique to optimize 14 key parameters related to GPU kernel scheduling, including the number of thread blocks and threads within a block. Our approach utilizes fully connected neural networks as the underlying machine learning model, with the tuning parameters as inputs to the neural networks and the actual execution time of a simulation as the outputs. To assess the effectiveness of our autotuning approach, we conducted experiments on three different types of GPUs, with computational speeds ranging from low to high. We performed independent training for each GPU model and also explored combined training across multiple GPU models. By leveraging artificial neural networks, our autotuning technique achieved remarkable results in tuning a wide range of parameters, leading to enhanced performance fora CFD code. Importantly, our approach demonstrated its efficacy while requiring only a small fraction of samples from the large parameter search space. This efficiency is attributed to the effectiveness of the fully connected neural networks in capturing the complex relationships between the parameter settings and the resulting performance. Overall, our study showcases the potential of machine learning, specifically fully connected neural networks, in autotuning GPU-accelerated CFD codes. By leveraging this approach, researchers and practitioners can achieve high performance in scientific simulations with optimized parameter configurations.  \n1 Introduction  \nGraphics processing units (GPUs), as highlighted in the study by Hwu et al.(1), have garnered significant attention in the field of scientific computing due to their enhanced computing capabilities and higher memory throughput in comparison to central processing units (CPUs) . CPUs typically serve as hosts that handle general settings and controls, while GPUs act as accelerator devices that execute intensive computations to achieve speedups. The GPU’s architecture, featuring thousands of lightweight cores, enables faster and more parallel computation synchronously on the device. Once  \nthe device completes the computations, the results are transferred back to the host. Consequently, data movements occur between the host and the device due to their distinct memories. The host and device can be interconnected through PCI-E or NVLink(2), which evidently enhances the memory bandwidth.  \nGPU exhibits multiple levels of parallelism, including block-level and thread-level parallelism, the tuning parameters of which are the block size k, the worker size m, and the vector length n in this example, as illustrated in Fig. 1. In GPU programming, the execution of a kernel function is orchestrated by a grid, which comprises a set of blocks or gangs. Each block consists of a group of threads that execute concurrently on the GPU’s streaming multiprocessors (SMs) . Within this parallel execution model, a worker represents the smallest unit of execution on the GPU, typically referring to individual threads or processing elements responsible for carrying out specific computations. Threads within a warp, a fundamental unit of execution in GPU architecture, are managed and scheduled together by the GPU’s hardware, enabling efficient execution of instructions. Collectively, these thread","cbCaivsoWUwpToCL","https://ap.wps.com/l/cbCaivsoWUwpToCL","pdf",759392,1,16,"English","en",105,"# Introduction\n## GPU acceleration and parallelism\n## CFD code and OpenACC background\n## Autotuning motivation and challenges","[{\"question\":\"What is the main goal of the proposed method?\",\"answer\":\"To optimize the performance of GPU-accelerated CFD applications by automatically tuning key kernel scheduling parameters.\"},{\"question\":\"Which tuning parameters are optimized in this study?\",\"answer\":\"The method optimizes 14 GPU kernel scheduling parameters, including the number of thread blocks and the threads within a block.\"},{\"question\":\"How does the machine learning model relate parameters to performance?\",\"answer\":\"Fully connected neural networks take tuning parameters as inputs and output the simulation execution time, enabling efficient prediction and selection.\"}]","Machine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance | PDF",1785727526,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-driven-autotuning-of-graphics-processing-unit-accelerated-computational-fluid-dynamics-for-enhanced-performance","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-driven-autotuning-of-graphics-processing-unit-accelerated-computational-fluid-dynamics-for-enhanced-performance/119991/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the proposed method?","Question",{"text":75,"@type":76},"To optimize the performance of GPU-accelerated CFD applications by automatically tuning key kernel scheduling parameters.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which tuning parameters are optimized in this study?",{"text":80,"@type":76},"The method optimizes 14 GPU kernel scheduling parameters, including the number of thread blocks and the threads within a block.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the machine learning model relate parameters to performance?",{"text":84,"@type":76},"Fully connected neural networks take tuning parameters as inputs and output the simulation execution time, enabling efficient prediction and selection.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]