[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81833-en":3,"doc-seo-81833-105":31,"detail-sidebar-cat-0-en-105":85},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81833,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Approximate Attention Weighting for Sustainable FPGA-Based Vision Transformer Inference","Vision Transformers use self-attention to capture global image context, but deploying them on small FPGAs is difficult because softmax requires exponentiation, summation, and normalization, which are expensive in hardware. This work introduces a BRAM-free approximate attention-weighting unit for FPGA-based ViT inference. The design uses a 16-segment piecewise-linear approximation of the natural exponential in softmax, fully implemented with distributed LUT fabric. On a Xilinx Zynq-7020, the attention-row core fits 1,444 LUTs and 77 DSPs with zero BRAM, achieving ≤0.20% absolute top-1 difference in hardware-accurate emulation while improving energy efficiency for edge AI.","arXiv :2607 .0 1798v2 [ cs .AR] 6 Jul 2026  \nAPPROXIMATE ATTENTION WEIGHTING FOR SUSTAINABLE FPGA-BASED VISION TRANSFORMER  \nINFERENCE  \nA PREPRINT  \nMuhammad Usman 1 , Muhammad Akmal Shafique2 , Shujaat Khan3 , and Dorit Merhof1  \n1Faculty of Informatics and Data Science, University of Regensburg, 93053 Regensburg, Germany  \n{muhammad.usman,[dorit.merhof](dorit.merhof}@ur.de)[}](dorit.merhof}@ur.de)[@ur.de](dorit.merhof}@ur.de)  \n2Department of Computer and Software Engineering, CEME, NUST, Pakistan  \n[akmal.shafique@yahoo.com](akmal.shafique@yahoo.com)  \n3Department of Computer Engineering, College of Computing and Mathematics, KFUPM,  \nDhahran, 31261, Saudi Arabia  \n[shujaat.khan@kfupm.edu.sa](shujaat.khan@kfupm.edu.sa)  \nABSTRACT  \nVision Transformers have reshaped computer vision by using self-attention to capture global context across image regions. This makes them attractive for edge visual inspection and monitoring in applications such as renewable-energy infrastructure, industrial quality control, medical imaging, and autonomous-system sensing. However, deploying ViTs on small FPGAs remains challenging because the softmax stage in self-attention requires exponential evaluation and normalization, which are costly in hardware. Existing implementations often rely on CORDIC pipelines or BRAM-based look-up tables, increasing area and power consumption. This paper presents a BRAM-free approximate attention-weighting unit for FPGA-based ViT inference. The proposed design approximates the natural exponential in softmax using a 16-segment piecewise-linear function implemented entirely with distributed LUT fabric. Unlike base-2 approximations, the natural-exponential formulation preserves the pre-trained attention temperature and avoids model-specific recalibration.  \nImplemented on a Xilinx Zynq-7020, the complete attention-row core uses 1444 LUTs, 77 DSPs, and no BRAM, while hardware-accurate emulation shows accuracy within a 0.20% absolute top-1 difference from the exact-softmax reference on ViT-family models. These results demonstrate the potential of the proposed core for energy-efficient ViT inference on resource-constrained edge-AI platforms.  \nKeywords Left-to-right arithmetic, FPGA, adder tree, ultrasound beamforming, dynamic precision, energy efficiency.  \n1 Introduction  \nVision Transformers (ViTs) [1, 2] have become widely used in computer vision because self-attention can model longrange spatial dependencies across image regions. This makes them effective for classification, detection, segmentation, and visual monitoring tasks in domains such as medical imaging, industrial quality control, smart-city sensing, autonomous-system perception, and infrastructure inspection [3] . As these workloads move from cloud servers to cameras, embedded controllers, agricultural sensors, and inspection nodes, inference energy becomes a continuous operational cost rather than a one-time training cost. Prior work has therefore argued that the energy and carbon footprint of AI should be treated as a design constraint, not as an afterthought [4, 5] .  \nThis concern is particularly relevant for sustainability-oriented edge intelligence, where large numbers of low-power devices may monitor renewable-energy infrastructure, smart grids, and industrial systems. ViT-based models have al-  \nUnder Review (IEEE 2026 Sustainable Energy and Industry Conference (SUSTAIN))  \nready been explored for photovoltaic-panel inspection [6], photovoltaic defect segmentation [7], wind-power forecasting [8], and electrical-load forecasting [9] . However, deploying ViTs on resource-limited FPGAs remains challenging because multi-head self-attention includes a softmax stage that requires exponentiation, summation, and division. While dot-product score computations map efficiently to FPGA DSP blocks, conventional softmax implementations often rely on CORDIC pipelines or BRAM-based look-up tables, increasing area, latency, and memory pressure. Base-2 approximatio","cbCaiqnqq3yPoEvS","https://ap.wps.com/l/cbCaiqnqq3yPoEvS","pdf",566383,5,1,10,"English","en",105,"# Abstract\n# Introduction\n## Problem: softmax on resource-limited FPGAs\n## Proposed BRAM-free attention-weighting unit","[{\"question\":\"What accuracy and resource results are reported on the Xilinx Zynq-7020?\",\"answer\":\"The complete attention-row core uses 1,444 LUTs, 77 DSPs, and zero BRAM, and hardware-accurate emulation shows accuracy within a 0.20% absolute top-1 difference from the exact-softmax reference on ViT-family models.\"}]","Approximate Attention Weighting for Sustainable FPGA-Based Vision Transformer Inference | PDF",1784176491,25,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":80,"head_meta":82,"extra_data":84,"updated_unix":29},"approximate-attention-weighting-for-sustainable-fpga-based-vision-transformer-inference","",{"@graph":37,"@context":79},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/approximate-attention-weighting-for-sustainable-fpga-based-vision-transformer-inference/81833/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73],{"name":74,"@type":75,"acceptedAnswer":76},"What accuracy and resource results are reported on the Xilinx Zynq-7020?","Question",{"text":77,"@type":78},"The complete attention-row core uses 1,444 LUTs, 77 DSPs, and zero BRAM, and hardware-accurate emulation shows accuracy within a 0.20% absolute top-1 difference from the exact-softmax reference on ViT-family models.","Answer","https://schema.org",{"og:url":53,"og:type":81,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":83,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":86},[87,91,95,99,103,108,113,116,121,124,127],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":88,"show_sort_weight":89,"slug":90},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":92,"show_sort_weight":93,"slug":94},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":47,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":47,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":22,"slug":126},"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":20,"slug":130},19,"General","general"]