[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82393-en":3,"doc-seo-82393-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82393,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Portable Acceleration of Learning With Errors KEMs for Post-Quantum Cryptography","The transition to post-quantum cryptography (PQC) demands real-world implementations meeting practical computational costs. Learning With Errors (LWE)-based key encapsulation mechanisms (KEMs) provide strong security foundations but incur heavy overhead from large matrix operations and extensive cryptographically secure random number generation. This work presents a portable GPU implementation of a plain LWE-based KEM using OpenMP Target offloading. By extending RNGonGPU with HIP support and integrating it into the OpenMP workflow, the solution achieves cross-platform acceleration on NVIDIA and AMD while avoiding vendor lock-in.","arXiv :2607 .0954 1v 1 [ cs .CR] 10 Jul 2026  \nPortable Acceleration of Learning With Errors KEMs for Post-Quantum Cryptography  \nTiziana Liberati 1 ,3 , Nitin Shukla2 , Simone Rizzo 1 , Elisabetta Boella 1 , Matteo Barbieri 1 ,4 ,5 , Gabriella  \nBettonte 1 , Daniele Gregori 1 , and Marco Pedicini3  \n1 E4 Computer Engineering SpA, via Martiri della Libertà 66, 42019 Scandiano (Italy) {tiziana.liberati,  \nsimone.rizzo, elisabetta.boella, matteo.barbieri, gabriella.bettonte, daniele.gregori}@e4company.com  \n2 SuperComputing Applications and Innovation Department, Cineca, via Magnanelli 6/3, 40033 Bologna (Italy) [n.shukla@cineca.it](n.shukla@cineca.it)  \n3 Dipartimento di Matematica e Fisica, Roma Tre University, Largo San Leonardo Murialdo 1, 00146 Rome (Italy) [marco.pedicini@uniroma3.it](marco.pedicini@uniroma3.it)  \n4 Dipartimento di Fisica e Astronomia “G. Galilei”, Università di Padova, via F. Marzolo 8, 35131 Padova (Italy)  \n5 INFN, Sezione di Padova, Italy  \nAbstract. The transition to post-quantum cryptography (PQC) is driving demand for implementations that can meet the computational requirements of real-world applications. Among the proposed PQC constructions, Learning With Errors (LWE) based key encapsulation mechanisms (KEMs) are particularly attractive due to their strong security foundations, but they incur substantial computational costs from matrix operations and large-scale cryptographically secure random number generation. These characteristics position GPU acceleration as an effective approach for lowering the computational overhead of lattice-based cryptographic schemes. In this work, we present a portable GPU implementation of a plain LWE-based KEM using OpenMP Target offloading. Unlike most existing GPU implementations, which rely on CUDA-specific optimizations, our approach uses a single source code base that executes on both NVIDIA and AMD accelerators. To enable fully GPU-resident execution on heterogeneous platforms, we extend the RNGonGPU library with the Heterogeneous-computing Interface for Portability (HIP) support and integrate it into the OpenMP Target workflow. We evaluate the proposed implementation on different accelerator architectures, analyzing performance benchmarking, runtime profiling, scalability analysis, and energy-to-solution measurements. Experimental results show that OpenMP Target offloading delivers substantial acceleration over a multicore CPU baseline while preserving source-level portability across heterogeneous GPU ecosystems. Cross-platform analysis identifies NVIDIA GH200 and AMD MI300X as the most effective platforms for this memory-bound workload, while profiling indicates that memory-system organization and CPU–GPU interaction play a more critical role than peak compute capability alone. These findings demonstrate that portable GPU acceleration can significantly reduce the computational overhead of PQC while avoiding vendor lock-in, thereby facilitating the deployment of quantum-resistant cryptographic infrastructures.  \nKeywords: Post-Quantum Cryptography · LWE-KEM · GPU Offloading · OpenMP · Performance Portability  \n1 Introduction  \nThe transition to post-quantum cryptography (PQC) has intensified the demand for implementations that can meet the computational requirements of real-world applications. In this work, we focus on lattice-based cryptography [10], one of the most widely studied families of post-quantum schemes, and in particular on Learning With Errors (LWE)-based key encapsulation mechanisms (KEMs) . The computational cost of these schemes is largely driven by large-scale matrix operations and large-scale  \n2 T. Liberati et al.  \ncryptographically secure random number generation [6] . Since these workloads expose a high degree of parallelism, GPUs are a natural platform for accelerating LWE-based cryptographic schemes. Previous GPU studies accelerated lattice-based PQC scheme such as FrodoKEM, NewHope, Kyber, and Dilithium [11,5] . These implementations sh","cbCaiqiKorY6KuxT","https://ap.wps.com/l/cbCaiqiKorY6KuxT","pdf",876767,1,15,"English","en",105,"# Abstract\n# Introduction\n## Motivation and Related Work\n## Contributions and Approach\n# Methodology and Evaluation","[{\"question\":\"Why are GPUs relevant for LWE-based KEM workloads in PQC?\",\"answer\":\"LWE-based KEMs are dominated by large matrix operations and large-scale cryptographically secure random number generation, which offer high parallelism. GPUs provide a natural platform to accelerate these components and reduce computational overhead.\"},{\"question\":\"What programming model and portability strategy does the work use?\",\"answer\":\"The implementation uses OpenMP Target offloading with a standards-based path intended to run across NVIDIA and AMD accelerators. It avoids CUDA-specific optimization paths to preserve source-level portability.\"},{\"question\":\"How is GPU-resident random number generation enabled across accelerator vendors?\",\"answer\":\"The work extends the RNGonGPU library by adding HIP support and integrates it into the OpenMP Target workflow. This enables fully GPU-resident execution on heterogeneous GPU platforms.\"}]",1784180103,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"portable-acceleration-of-learning-with-errors-kems-for-post-quantum-cryptography","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/portable-acceleration-of-learning-with-errors-kems-for-post-quantum-cryptography/82393/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are GPUs relevant for LWE-based KEM workloads in PQC?","Question",{"text":75,"@type":76},"LWE-based KEMs are dominated by large matrix operations and large-scale cryptographically secure random number generation, which offer high parallelism. GPUs provide a natural platform to accelerate these components and reduce computational overhead.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What programming model and portability strategy does the work use?",{"text":80,"@type":76},"The implementation uses OpenMP Target offloading with a standards-based path intended to run across NVIDIA and AMD accelerators. It avoids CUDA-specific optimization paths to preserve source-level portability.",{"name":82,"@type":73,"acceptedAnswer":83},"How is GPU-resident random number generation enabled across accelerator vendors?",{"text":84,"@type":76},"The work extends the RNGonGPU library by adding HIP support and integrates it into the OpenMP Target workflow. This enables fully GPU-resident execution on heterogeneous GPU platforms.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]