[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81776-en":3,"doc-seo-81776-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81776,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","High-Performance NTT Accelerators for PQC Leveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture","Post-quantum cryptography and privacy-preserving technologies require fast polynomial arithmetic, where the Number Theoretic Transform (NTT) is a core primitive for lattice-based schemes such as ML-KEM (CRYSTALS-Kyber) and ML-DSA (CRYSTALS-Dilithium). The document proposes parallel iterative NTT/INTT accelerators using unified butterfly hardware, a new redundant number representation to remove conditional corrections, and inverse-transform scaling integrated into existing arithmetic. Hierarchical Montgomery multipliers are mapped to FPGA DSP resources to lower cost and raise operating frequency. FPGA results show reduced execution time and improved clock rates with competitive resource usage.","arXiv :2607 .0062 1v 1 [ cs .AR] 1 Jul 2026  \nGraphical Abstract  \nHigh-Performance NTT Accelerators for PQC leveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture  \nGeorge Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos  \nHighlights  \nHigh-Performance NTT Accelerators for PQC leveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture  \nGeorge Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos  \n• Unified redundant arithmetic for NTT/INTT: Introduces a novel redundant number representation that extends Montgomery-based redundancy to support combined subtract–multiply operations, enabling a single butterfly unit to operate efficiently in both NTT and INTT modes while significantly reducing modulo correction overhead.  \n• Microarchitectural Fine Tuning : Merges INTT scaling operations into existing arithmetic hardware and hierarchically maps Montgomery multipliers onto FPGA DSP blocks, simplifying pipelining, lowering hardware cost, and enabling higher operating frequencies.  \n• Consistent performance gains across PQC-relevant scales: The proposed design achieves significant latency reductions over state-ofthe-art NTT accelerators for runtime-programmable moduli, while maintaining comparable or smaller resource utilization and remaining competitive with fixed-modulus designs.  \nHigh-Performance NTT Accelerators for PQCleveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture  \nGeorge Alexakisa , Dimitrios Schoinianakisb , Giorgos Dimitrakopoulosa  \na Electrical and Computer Engineering, Democritus University of Thrace, Xanthi, Greece  \nb Nokia Bell Labs, Athens, Greece  \nAbstract  \nPost-quantum cryptography and privacy-preserving technologies are expected to play a central role in future secure communication systems. Latticebased PQC schemes such as ML-KEM (CRYSTALS-Kyber) and ML-DSA (CRYSTALS-Dilithium) rely heavily on large-degree polynomial arithmetic, making the Number Theoretic Transform (NTT) a key computational primitive. Although existing hardware accelerators exploit parallelism and pipelining to support both NTT and INTT, their efficiency is often limited by the overhead of modular reduction and correction steps, inverse-transform scaling operations, and suboptimal FPGA implementations. This work addresses these limitations by proposing parallel iterative NTT/INTT accelerators based on optimized unified butterfly units. We introduce a novel redundant number representation that eliminates conditional corrections for both Montgomery modulo multiplication and combined subtract–multiply operations, and integrate inverse-transform scaling into existing arithmetic hardware to avoid dedicated scaling units. Furthermore, we design hierarchical Montgomery multipliers that map efficiently onto FPGA DSP resources, reducing hardware cost while enabling high operating frequencies. FPGAbased experimental results demonstrate higher clock frequencies, reduced execution times, and competitive resource utilization, supporting efficient NTT acceleration for PQC and related privacy-preserving applications.  \nKeywords: Number theoretic transform, Post Quantum Cryptography, Hardware Accelerators  \n1. Introduction  \nPost-Quantum Cryptography (PQC) and Privacy-Preserving Technologies (PPT) stand as the modern pillars of long-term secure communications, integrity, and data monetization. Their integration is particularly vital for the upcoming 6G era, where PQC is set to be a native security requirement from the beginning to defend against the threats posed by quantum computers [1] . Future quantum computers will easily break the public-key algorithms that currently protect global digital infrastructure and rely on conventionally hard mathematical problems, such as integer factorization and the discrete logarithm, thereby creating an immediate security vulnerability [2] . At the same time, 6G networks are expected to leverage advanced PPTs, most notably Fully Homomorphic Encrypti","cbCaifQ88skL9TIq","https://ap.wps.com/l/cbCaifQ88skL9TIq","pdf",3551494,4,1,37,"English","en",105,"# Highlights\n## Unified redundant arithmetic for NTT/INTT\n## Microarchitectural fine tuning\n## Performance gains across PQC-relevant scales\n# Abstract\n# 1. Introduction","[{\"question\":\"What problem does the proposed design address in existing NTT/INTT FPGA accelerators?\",\"answer\":\"It targets overhead from modular reduction/correction steps, inverse-transform scaling operations, and inefficient FPGA implementations that limit accelerator efficiency.\"},{\"question\":\"How does unified redundant arithmetic improve NTT and INTT execution?\",\"answer\":\"It introduces a redundant number representation extending Montgomery-based redundancy so a single butterfly unit can handle both NTT and INTT efficiently while reducing modulo correction overhead and eliminating conditional corrections for key operations.\"},{\"question\":\"What microarchitecture techniques are used to improve FPGA efficiency?\",\"answer\":\"Inverse-transform scaling is merged into existing arithmetic hardware, and hierarchical Montgomery multipliers are mapped onto FPGA DSP blocks to simplify pipelining, reduce hardware cost, and enable higher operating frequencies.\"}]",1784176083,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"high-performance-ntt-accelerators-for-pqc-leveraging-unified-redundant-arithmetic-and-fine-tuned-microarchitecture","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/high-performance-ntt-accelerators-for-pqc-leveraging-unified-redundant-arithmetic-and-fine-tuned-microarchitecture/81776/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the proposed design address in existing NTT/INTT FPGA accelerators?","Question",{"text":75,"@type":76},"It targets overhead from modular reduction/correction steps, inverse-transform scaling operations, and inefficient FPGA implementations that limit accelerator efficiency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does unified redundant arithmetic improve NTT and INTT execution?",{"text":80,"@type":76},"It introduces a redundant number representation extending Montgomery-based redundancy so a single butterfly unit can handle both NTT and INTT efficiently while reducing modulo correction overhead and eliminating conditional corrections for key operations.",{"name":82,"@type":73,"acceptedAnswer":83},"What microarchitecture techniques are used to improve FPGA efficiency?",{"text":84,"@type":76},"Inverse-transform scaling is merged into existing arithmetic hardware, and hierarchical Montgomery multipliers are mapped onto FPGA DSP blocks to simplify pipelining, reduce hardware cost, and enable higher operating frequencies.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]