[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81941-en":3,"doc-seo-81941-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81941,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Boosting FPGA Performance with Direct BRAM-DSP Paths","Efficient memory-to-compute data movement is a dominant performance bottleneck in modern FPGA designs, especially for deep learning inference. In conventional architectures, BRAM-to-DSP transfers traverse global routing, increasing wirelength, routing congestion, and critical-path delays. The paper introduces a lightweight architectural enhancement with dedicated direct BRAM–DSP connections plus a placement-aware CAD update. The change preserves backward compatibility and adds negligible area/timing overhead, improving Agilex-10-like DL layer designs by up to +25% Fmax and −49% wirelength.","Boosting FPGA Performance with Direct  \nBRAM-DSP Paths  \nJiajun Hu 1 , Ruthwik Reddy Sunketa 1 , Andrew Boutros2 , Aman Arora 1  \n1Arizona State University, Tempe, AZ, USA 2University of Waterloo, Waterloo, Ontario, CA  \n{jiajunh5, rsunketa, [aman.kbm](aman.kbm}@asu.edu)[}](aman.kbm}@asu.edu)[@asu.edu](aman.kbm}@asu.edu) [andrew.boutros@uwaterloo.ca](andrew.boutros@uwaterloo.ca)  \narXiv :2607 .05756v 1 [ cs .AR] 7 Jul 2026  \nAbstract—Efficient data movement between memory and compute units is a key performance bottleneck in modern FPGA designs, particularly for deep learning (DL) workloads. In typical FPGA architectures, data transfers between block RAMs (BRAMs) and digital signal processing units (DSPs) must traverse the global routing network, leading to increased wirelength, routing congestion, and critical-path delays. Prior work has explored in- and near-BRAM compute architectures to mitigate these issues, but such solutions often require fundamental changes to FPGA architecture and CAD tools, limiting their commercial viability. This paper proposes a lightweight architectural enhancement that introduces a dedicated direct connection between BRAM and DSP blocks, enabling BRAM data to be consumed by DSPs without passing through the global interconnect. We also enhance the placement algorithm to recognize these BRAM–DSP macro blocks. The proposed architectural change incurs negligible area and delay overhead and does not affect non-DL benchmarks, while the proposed CAD remains compatible with the baseline architecture, where it yields negligible change in quality-of-results (QoR). On an Agilex-10-like FPGA, the proposed architecture and CAD updates deliver up to +25% Fmax and −49% wirelength on common DL layer designs.  \nI. INTRODUCTION  \nAs modern compute workloads scale significantly, data movement is a growing bottleneck. This impacts FPGAsas well: data must frequently move between various blocks through programmable routing. A characteristic feature of contemporary FPGA architectures is the inclusion of dedicated direct interconnects between homogeneous blocks within a column, which can be exploited to reduce reliance on the programmable routing network thereby reducing data movement. These include cascading BRAMs into deeper logical memories, chaining DSPs for MAC operations and CLB carry chains for larger adders [1],[2] . However, these hardened paths do not address data movement between heterogeneous blocks, particularly between BRAMs and DSPs, which still traverses the global routing network.  \nThis limitation is acute in Deep Learning (DL) inference, which is dominated by General Matrix–Vector Multiplication (GEMV) . In a GEMV operation, a DSP multiplies a BRAMresident weight against a runtime-streamed activation through its two input ports. Because DL accelerators [3] instantiate thousands of such BRAM-DSP pairs operating in parallel, the resulting high-volume BRAM-to-DSP traffic has no direct path, so it traverses the global routing network, creating routing congestion and lengthening the critical path. Recent in-/near-BRAM compute architectures cut this data movement but require fundamental changes to the BRAM architecture  \n& circuitry as well as the CAD toolchain, raising area, programming, and commercial-viability concerns [4], [5] .  \nTo enable efficient BRAM-to-DSP data transfer without disrupting existing FPGA development flows, we propose an enhanced FPGA architecture that has direct paths between BRAMs and DSPs. BRAM output is directly connected to the input of a DSP a few columns away, directly facing this BRAM. Importantly, the original BRAM and DSP routing interfaces are preserved, maintaining full backward compatibility. Since current CAD flows support only homogeneous (same-type) direct interconnects, they cannot correctly leverage the proposed heterogeneous (cross-type) connections, leading to placement and routing failures. We enhance the placement engine with simple updates to fully utilize the p","cbCaitb9DNWK4qKc","https://ap.wps.com/l/cbCaitb9DNWK4qKc","pdf",416461,5,1,4,"English","en",105,"# I. Introduction\n# II. Background & Related Work","[{\"question\":\"What performance issue does the paper target in FPGA deep learning workloads?\",\"answer\":\"BRAM-to-DSP data transfers must use the global routing network, which increases wirelength, routing congestion, and critical-path delays and limits achievable clock frequency.\"},{\"question\":\"What architectural enhancement is proposed to improve BRAM-to-DSP communication?\",\"answer\":\"A hardened, cross-type direct connection between BRAM output ports and DSP inputs is added so BRAM data can feed DSPs without passing through global interconnects.\"},{\"question\":\"How does the proposed CAD flow support the new BRAM–DSP direct connections?\",\"answer\":\"The placement engine is extended to recognize cross-type BRAM–DSP direct links as placement macros and to update placement priorities after each successful placement.\"}]","Boosting FPGA Performance with Direct BRAM-DSP Paths | PDF",1784177180,10,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"boosting-fpga-performance-with-direct-bram-dsp-paths","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":22},"https://docshare.wps.com/document/boosting-fpga-performance-with-direct-bram-dsp-paths/81941/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What performance issue does the paper target in FPGA deep learning workloads?","Question",{"text":76,"@type":77},"BRAM-to-DSP data transfers must use the global routing network, which increases wirelength, routing congestion, and critical-path delays and limits achievable clock frequency.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What architectural enhancement is proposed to improve BRAM-to-DSP communication?",{"text":81,"@type":77},"A hardened, cross-type direct connection between BRAM output ports and DSP inputs is added so BRAM data can feed DSPs without passing through global interconnects.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed CAD flow support the new BRAM–DSP direct connections?",{"text":85,"@type":77},"The placement engine is extended to recognize cross-type BRAM–DSP direct links as placement macros and to update placement priorities after each successful placement.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":30,"doc_module":4,"doc_module_name":47,"category_name":132,"show_sort_weight":30,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":47,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]