[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82028-en":3,"doc-seo-82028-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82028,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","FPGN Redefining Ultra Fast Programmable Gate based Neural Acceleration with Differentiable LUTs","Deep neural network inference at nanosecond scale is a core architectural constraint for latency-critical systems, yet FPGA accelerators typically remain arithmetic-centric, using LUTs mainly for numerical operators rather than exploiting LUTs as intrinsically expressive logic. Prior differentiable LUT-native approaches fail to faithfully model FPGA LUT primitives, use physically-unaware topologies that hurt routability and timing closure, and lack automated optimization flows. This paper introduces FPGN, a physically aware end-to-end framework that aligns differentiable LUT training, improves streaming topology for timing, and adds a latency-driven compiler for automated DSE and hardware generation.","FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs  \nJiawei Liang 1 , Haotong Qin2 , Linfeng Du 1 , Xingyu Liu 1 , Shangkun Li 1 , Hui Yu 1 , Michele Magno2 , Xinyu Chen3 , Jiang Xu3 , Wei Zhang 1  \n1HKUST, 2ETH Zurich, 3HKUST (GZ)  \narXiv :2607 .08427v 1 [ cs .AR] 9 Jul 2026  \nAbstract—Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. While FieldProgrammable Gate Arrays (FPGAs) offer a promising substrate for low-latency inference, conventional FPGA accelerators remain arithmetic-centric, using LUTs primarily as building blocks for numerical operators and peripheral logic. In contrast, recent LUT-native neural networks treat LUTs as learnable neurons, revealing promising theoretical potential to exploit their intrinsic logic expressivity. However, existing methods are largely confined to algorithmic optimizations, failing to translate this theoretical potential into high-performance FPGA accelerators. Specifically, their differentiable formulations do not faithfully match FPGALUT primitives, their physically-unaware topologies compromise routability and timing closure, and their lack of automated optimization flow hinders systematic design space exploration (DSE) and efficient hardware implementation.  \nIn this paper, we propose FPGN, an end-to-end physicallyaware framework that closes the gap between LUT-native learning and latency-optimized FPGA implementation. FPGN addresses these challenges through (i) a hardware-aligned differentiable formulation for training FPGA-native LUT neurons,(ii) a structured LUT-native topology with a streaming hardware architecture to improve routing locality and timing closure, and (iii) a latency-driven compiler that leverages high-fidelity analytical Quality of Results models to automate DSE and hardware generation. Experiments show that FPGN achieves up to 205× latency reduction compared to representative FPGAbased BNN accelerators and up to 30 × higher LUT efficiency than prior differentiable LUT-native networks, while maintaining competitive inference accuracy.  \nIndex Terms—Differentiable, FPGA, LUT-Native Networks, Hardware Co-Design.  \nI. INTRODUCTION  \nDeep Neural Networks (DNNs) have driven major advancesin artificial intelligence [1], [2] . However, deploying them in domains prioritizing nanosecond-scale response requirements over other constraints remains highly challenging, such as high-energy physics triggers [3], [4], high-frequency trading [5], [6], and line-rate packet or flow classification in highspeed network data planes [7], [8] . While GPUs dominate high-throughput training, their rigid memory hierarchy and batch-oriented execution make them architecturally unsuitable for these latency-critical applications [9] .  \nSpecialized accelerators on Field-Programmable Gate Arrays (FPGAs) [10], [11] and Application-Specific Integrated Circuits (ASICs) [12], [13] have been developed to reduce inference latency. While ASICs offer peak performance, their inflexibility and high Non-Recurring Engineering (NRE) costs limit their adaptability to rapidly evolving algorithms [14] .  \nFig. 1. Evolution of neural paradigms on FPGA. (a) MAC in CNN. (b) BNNs replace MAC with XNOR and N-input popcount. (c) The LUT-asoperator paradigm replaces XNOR with learnable k-LUTs while maintaining arithmetic-centric structures. (d) LUT-as-neuron paradigm unifies operators and weights into learnable LUTs.  \nConsequently, FPGAs have emerged as a compelling substrate for hardware-algorithm co-design, offering agile reconfigurability alongside deterministic ultra-low latency. At the core of this flexibility is the k-input Look-Up Table (k-LUT), a small memory that can be programmed to realize any Boolean function of k input bits.  \nDespite inherent programmability, traditional FPGA accelerators follow an arithmetic-centric paradigm that treats LUTs as building","cbCaihKCfOAdYZcR","https://ap.wps.com/l/cbCaihKCfOAdYZcR","pdf",4466939,7,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does FPGN target in FPGA-based neural acceleration?\",\"answer\":\"FPGN targets achieving nanosecond-scale inference latency while overcoming the gap between LUT-native learning theory and latency-optimized FPGA implementation.\"},{\"question\":\"Why do existing differentiable LUT-native methods underperform on real FPGA hardware?\",\"answer\":\"They do not accurately match FPGA LUT primitives, use physically-unaware topologies that compromise routability and timing closure, and lack automated optimization flow for efficient hardware generation.\"},{\"question\":\"How does FPGN improve training and hardware implementation of LUT-native LUT neurons?\",\"answer\":\"It uses a hardware-aligned differentiable formulation for training FPGA-native LUT neurons, a structured LUT-native streaming topology to improve routing locality and timing, and a latency-driven compiler with high-fidelity analytical QoR models to automate DSE and hardware generation.\"}]","FPGN Redefining Ultra Fast Programmable Gate based Neural Acceleration with Differentiable LUTs | PDF",1784177683,35,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"fpgn-redefining-ultra-fast-programmable-gate-based-neural-acceleration-with-differentiable-luts","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/fpgn-redefining-ultra-fast-programmable-gate-based-neural-acceleration-with-differentiable-luts/82028/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does FPGN target in FPGA-based neural acceleration?","Question",{"text":77,"@type":78},"FPGN targets achieving nanosecond-scale inference latency while overcoming the gap between LUT-native learning theory and latency-optimized FPGA implementation.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Why do existing differentiable LUT-native methods underperform on real FPGA hardware?",{"text":82,"@type":78},"They do not accurately match FPGA LUT primitives, use physically-unaware topologies that compromise routability and timing closure, and lack automated optimization flow for efficient hardware generation.",{"name":84,"@type":75,"acceptedAnswer":85},"How does FPGN improve training and hardware implementation of LUT-native LUT neurons?",{"text":86,"@type":78},"It uses a hardware-aligned differentiable formulation for training FPGA-native LUT neurons, a structured LUT-native streaming topology to improve routing locality and timing, and a latency-driven compiler with high-fidelity analytical QoR models to automate DSE and hardware generation.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]