[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-148802-105":59,"doc-detail-148802-en":134},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":127,"head_meta":129,"extra_data":131,"updated_unix":133},105,"en","stencil-computations-on-amd-and-nvidia-graphics-processors-performance-and-tuning-strategies","Stencil Computations on AMD and Nvidia Graphics Processors - Performance and Tuning Strategies","","Over the last ten years, graphics processors have become the default accelerators for data-parallel workloads in high-performance computing, including machine learning and computational sciences. With AMD GPUs now deployed on leading supercomputers, established tuning methods for older hardware require reassessment. This study measures performance and energy efficiency of stencil computations on modern datacenter GPUs and introduces a cache-fusion tuning strategy for cache-heavy stencil kernels. Experiments cover linear and nonlinear stencils in one to three dimensions using both synthetic and practical applications, showing AMD and Nvidia differences that require platform-specific tuning.",{"@graph":69,"@context":126},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/stencil-computations-on-amd-and-nvidia-graphics-processors-performance-and-tuning-strategies/148802/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/stencil-computations-on-amd-and-nvidia-graphics-processors-performance-and-tuning-strategies/148802.png","ImageObject",300,407,{"name":92,"@type":93},"Clementine","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-17","2026-08-26",true,{"@type":102,"interactionType":103,"userInteractionCount":24},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118,122],{"name":109,"@type":110,"acceptedAnswer":111},"为什么需要重新评估针对GPU的调优策略？","Question",{"text":112,"@type":113},"随着AMD制造的GPU进入世界最快超算，旧硬件世代的调优方法不再完全适用，需要针对新平台重新评估。","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"本文研究的调优策略是什么？",{"text":117,"@type":113},"提出用于“缓存密集型”stencil内核的调优方法，核心是对相关计算进行fusing以提升缓存利用。",{"name":119,"@type":110,"acceptedAnswer":120},"实验覆盖了哪些类型的stencil计算？",{"text":121,"@type":113},"实验同时包含合成与实际应用，评估一维到三维的线性与非线性stencil函数的表现。",{"name":123,"@type":110,"acceptedAnswer":124},"AMD和Nvidia GPU在本文中体现出哪些关键差异？",{"text":125,"@type":113},"作者指出，在硬件与软件层面都存在重要差异，因此需要进行平台特定调优才能释放各自的计算潜力。","https://schema.org",{"og:url":83,"og:type":128,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":130,"canonical":83},"index,follow",{"doc_id":132,"site_id":62},148802,1787786236,{"code":4,"msg":5,"data":135},{"doc_id":132,"user_id":136,"nickname":92,"user_avatar":137,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":138,"file_id":139,"file_url":140,"file_type":141,"file_size":142,"view_count":24,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":143,"language":144,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":145,"faqs":146,"seo_title":147,"seo_description":67,"update_tm":133,"read_time":26},1374391974564,"https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002","This is an electronic reprint of the original article.  \nThis reprint may differ from the original in pagination and typographic detail.  \nPekkilä, Johannes; Lappi, Oskar; Robertsen, Fredrik; Korpi-Lagg, Maarit  \nStencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies  \nPublished in:  \nConcurrency and Computation: Practice and Experience  \nDOI:  \n10.1002/cpe.70129  \nPublished: 25/06/2025  \nDocument Version  \nPublisher's PDF, also known as Version of record  \nPublished under the following license:  \nCC BY  \nPlease cite the original version:  \nPekkilä, J. , Lappi, O. , Robertsen, F. , & Korpi-Lagg, M. (2025) . Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies. Concurrency and Computation: Practice and  \nExperience, 37(12-14), 1-23 . Article e70129 . [https://doi.org/10.1002/cpe.70129](https://doi.org/10.1002/cpe.70129)  \nThis material is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you foryour research use or educational purposes in electronic or print form. You must obtain permission for anyother use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user.  \nConcurrency and Computation: Practice and Experience  \nRESEARCH ARTICLE  OPEN ACCESS   \nStencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies  \nJohannes Pekkilä1  | Oskar Lappi2  | Fredrik Robertsén3 | Maarit J. Korpi-Lagg1,4,5  \n1 Department of Computer Science, Aalto University, Espoo, Finland | 2 Department of Computer Science, University of Helsinki, Helsinki, Finland |  \n3 CSC—IT Center for Science Ltd, Espoo, Finland | 4Max Planck Institute for Solar System Research, Göttingen, Germany | 5Nordita, KTH Royal Institute of Technology and Stockholm University, Stockholm, Sweden  \nCorrespondence: Johannes Pekkilä ([johannes.pekkila@aalto.fi](johannes.pekkila@aalto.fi))  \nReceived: 5 June 2024 | Revised: 11 September 2024 | Accepted: 7 May 2025  \nFunding: This work was supported by the Academy of Finland, ReSoLVE Centre of Excellence (grant number 307411), the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (Project UniSDyn, grant agreement no: 818665), and KAUTE foundation (grant number 20240173) .  \nKeywords: discrete convolution | energy efficiency | graphics processing units | high-performance computing | partial differential equations | performance optimization | stencil computations  \nABSTRACT  \nOver the last ten years, graphics processors have become the de facto accelerator for data-parallel tasks in various branches of high-performance computing, including machine learning and computational sciences. However, with the recent introduction of AMD-manufactured graphics processors to theworld’s fastest supercomputers, tuning strategies established for previous hardware generations must be re-evaluated. In this study, we evaluate the performance and energy efficiency of stencil computations on modern datacenter graphics processors and propose a tuning strategy for fusing cache-heavy stencil kernels. The studied cases comprise both synthetic and practical applications, which involve the evaluation of linear and nonlinear stencil functions in one to three dimensions. Our experiments reveal that AMD and Nvidia graphics processors exhibit key differences in both hardware and software, necessitating platform-specific tuning to reach their full computational potential.  \n1 | Introduction  \nStencil computations belong to a class of algorithms, where the elements of an array are updated by extracting information from their neighborhood in a fixed pattern, called a stencil (Figure 1) . A typical example is median filtering in image processing, where the color of each pixel is set to the med","cbCairpLjTXfek6K","https://ap.wps.com/l/cbCairpLjTXfek6K","pdf",4475894,24,"English","# Introduction\n## Stencil computation background and examples\n## GPUs in HPC and motivation for retuning\n# Methods and tuning strategy\n## Fusing cache-heavy stencil kernels\n## Studied linear and nonlinear stencils\n# Experiments and results\n## Performance and energy efficiency comparison\n## AMD vs Nvidia hardware/software differences\n# Conclusion","[{\"question\":\"为什么需要重新评估针对GPU的调优策略？\",\"answer\":\"随着AMD制造的GPU进入世界最快超算，旧硬件世代的调优方法不再完全适用，需要针对新平台重新评估。\"},{\"question\":\"本文研究的调优策略是什么？\",\"answer\":\"提出用于“缓存密集型”stencil内核的调优方法，核心是对相关计算进行fusing以提升缓存利用。\"},{\"question\":\"实验覆盖了哪些类型的stencil计算？\",\"answer\":\"实验同时包含合成与实际应用，评估一维到三维的线性与非线性stencil函数的表现。\"},{\"question\":\"AMD和Nvidia GPU在本文中体现出哪些关键差异？\",\"answer\":\"作者指出，在硬件与软件层面都存在重要差异，因此需要进行平台特定调优才能释放各自的计算潜力。\"}]","Stencil Computations on AMD and Nvidia Graphics Processors - Performance and Tuning Strategies | PDF"]