[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84707-en":3,"doc-seo-84707-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84707,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Kaizen: Metamorphic Fuzzing and Differential Testing for LLM-Translated HPC Applications","Large language models increasingly translate scientific codes across heterogeneous high-performance computing (HPC) programming models, such as CUDA to OpenMP, OpenACC, Kokkos, or SYCL. Existing evaluations often rely on compilation success, token-level similarity, or developer-written tests, which do not reliably guarantee behavioral correctness. Kaizen introduces a metamorphic fuzzing and differential testing framework to generate semantically equivalent mutants, diversify inputs, and detect semantic divergences when translated programs compile and pass tests but yield incorrect scientific results.","arXiv :2607 .04058v 1 [ cs . SE] 4 Jul 2026  \nKaizen: Metamorphic Fuzzing and Differential Testing for LLM-Translated HPC Applications  \nOSCAR LUDWIG, Oregon State University, USA  \nNINAD ANKLESARIA, Oregon State University, USA ZHEMING JIN∗ , Oak Ridge National Laboratory, USASWAROOP POPHALE, Oak Ridge National Laboratory, USA KAUSAR MOSHOOD†, Oregon State University, USA CHRISTIAN J. DEVORE†, Oregon State University, USA BRANDON GILL†, Oregon State University, USA CASSIUS VILLAREAL†, Oregon State University, USA KEITA TERANISHI, Oak Ridge National Laboratory, USAMANISH MOTWANI✉ , Oregon State University, USA  \nLarge language models (LLMs) are increasingly used to port scientific codes across heterogeneous high-performance computing (HPC) programming models, such as translating CUDA to OpenMP, OpenACC, Kokkos or SYCL. However, current evaluationsuse compilation success, token-level similarity, or developer-written tests from static benchmarks, which cannot reliably ensure behavioral correctness. We present Kaizen, a metamorphic fuzzing and differential testing framework for evaluating the correctness of LLM-translated HPC code. Kaizen uses metamorphic fuzzing via source-code mutation to generate semantically equivalent programs, grammar-based input fuzzing to explore behavioral diversity, and differential testing to expose semantic divergences between original and translated applications that compile and pass developer-written tests yet produce incorrect scientific results.  \nWe evaluate Kaizen on CUDA-to-OpenMP translation of 16 scientific applications from seven domains using three fine-tuned LLMs at kernel-level and full-program granularity. Our evaluation reveals that (1) compilation success is a poor proxy for correctness; (2) LLM-translated programs exhibit systematic compile-time error patterns, with nine categories for kernel-level translation and 27 for full-program translation; (3) semantic errors that survive compilation are often input-dependent and require differential testing to expose; and (4) full-program translation is substantially harder than kernel-level translation. These findings highlight the need for correctness-oriented evaluation of LLM-assisted HPC code translations.  \nCCS Concepts: • Software and its engineering → Software testing and debugging; Software verification and validation; • Computing methodologies → Machine learning; Artificial intelligence.  \n1 Introduction  \nAfter the demise of Moore’s law and Dennard scaling, modern high-performance computing (HPC) platforms have evolved primarily by increasing parallelism through multicore CPUs, then manycore architectures, and now accelerator (GPUs) based systems. These architectural shifts have driven continual evolution of programming models and their  \n∗ Work was completed while affiliated with Oak Ridge National Laboratory †Authors contributed equally to this research.  \nAuthors’ Contact Information: Oscar Ludwig, Oregon State University, Corvallis, Oregon, USA, [ludwigo@oregonstate.edu](ludwigo@oregonstate.edu); Ninad Anklesaria, Oregon State  \nUniversity, Corvallis, Oregon, USA, [anklesan@oregonstate.edu](anklesan@oregonstate.edu); Zheming Jin, Oak Ridge National Laboratory, Oak Ridge, USA, [zheming.jin@gmail.com](zheming.jin@gmail.com);  \nSwaroop Pophale, Oak Ridge National Laboratory, Oak Ridge, USA, [pophaless@ornl.gov](pophaless@ornl.gov); Kausar Moshood, Oregon State University, Corvallis, Ore  \ngon, USA, [moshoodk@oregonstate.edu](moshoodk@oregonstate.edu); Christian J. DeVore, Oregon State University, Corvallis, Oregon, USA, [devorech@oregonstate.edu](devorech@oregonstate.edu); Brandon  \nGill, Oregon State University, Corvallis, Oregon, USA, [gillb3@oregonstate.edu](gillb3@oregonstate.edu); Cassius Villareal, Oregon State University, Corvallis, Oregon, USA, [villarec@oregonstate.edu](villarec@oregonstate.edu); Keita Teranishi, Oak Ridge National Laboratory, Oak Ridge, USA, [teranishik@ornl.gov](teranishik@ornl.gov); Manish Motwani, Oreg","cbCaiilAyfz3sYQt","https://ap.wps.com/l/cbCaiilAyfz3sYQt","pdf",847561,1,28,"English","en",105,"# Introduction\n## Kaizen framework overview\n# Evaluation results\n## Error patterns in kernel and full-program translation","[{\"question\":\"What key findings emerge from evaluating Kaizen on CUDA-to-OpenMP translations?\",\"answer\":\"Compilation success is a poor correctness proxy, compile-time errors follow systematic patterns across kernel-level and full-program translations, many semantic errors depend on inputs and require differential testing, and full-program translation is harder than kernel-level translation.\"}]",1784197763,71,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"kaizen-metamorphic-fuzzing-and-differential-testing-for-llm-translated-hpc-applications","",{"@graph":35,"@context":76},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/kaizen-metamorphic-fuzzing-and-differential-testing-for-llm-translated-hpc-applications/84707/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What key findings emerge from evaluating Kaizen on CUDA-to-OpenMP translations?","Question",{"text":74,"@type":75},"Compilation success is a poor correctness proxy, compile-time errors follow systematic patterns across kernel-level and full-program translations, many semantic errors depend on inputs and require differential testing, and full-program translation is harder than kernel-level translation.","Answer","https://schema.org",{"og:url":51,"og:type":78,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":80,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general"]