[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81900-en":3,"doc-seo-81900-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81900,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","A Reconfigurable and Representation Adaptive ISA Based Architecture for Efficient DNN Acceleration","Machine-learning-oriented deep neural network (DNN) acceleration is presented as a solution to a core tradeoff: domain-specific accelerators deliver high performance and energy efficiency but remain inflexible to evolving model architectures, while ISA-based approaches improve programmability at an efficiency cost. The proposed ISA and reconfigurable platform provide fine-grained control over data movement, dynamic precision, and decoupled execution. A Residue Number System instantiation with 3–8-bit dynamic precision is implemented in 22 nm, reaching 5.12–10.47 TOPS/W and up to 1.2× higher energy efficiency than fixed-point while preserving model accuracy.","A Reconfigurable and Representation-Adaptive ISA-Based Architecture for Efficient DNN  \nAcceleration  \nVasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis  \narXiv :2607 .04475v 1 [ cs .AR] 5 Jul 2026  \nAbstract—Domain-specific hardware accelerators provide significantly higher performance and energy efficiency for deep neural network (DNN) workloads than general-purpose processors, but often lack adaptability to evolving model architectures. In contrast, general-purpose ISA-based solutions, such as RISCV-based accelerators, improve programmability at the cost of efficiency. This work addresses this tradeoff by introducing a machine-learning-oriented instruction set architecture (ISA) anda reconfigurable hardware platform that combine high efficiency with flexibility. The proposed ISA enables fine-grained control over data movement, dynamic precision, and decoupled execution across data-fetching, tensor processing, and post-processing domains. The corresponding architecture employs lightweight programmable cores and SIMD units to maintain high processingelement utilization with low control overhead, while remaining independent of the underlying numerical representation. We demonstrate the approach using a Residue Number System (RNS) instantiation supporting 3–8-bit dynamic precision. A 22-nm implementation achieves 5.12–10.47 TOPS/W for a typical workload and up to 1.2 × higher energy efficiency than its fixed-point counterpart, while preserving model accuracy. It also outperforms state-of-the-art and mixed-precision accelerators. These results show that the proposed design effectively bridges the gap between efficiency and programmability in modern DNN accelerators.  \nIndex Terms—AI accelerator, ISA, RNS  \nI. INTRODUCTION  \nRecent advances in domain-specific hardware accelerators have delivered substantial performance and efficiency gains for AI workloads, particularly in computer vision (e.g., CNNs) and natural language processing (e.g., transformers) . Stateof-the-art AI accelerators can achieve power efficiencies of 10 − 50 TOPS/W [1], [2], [3], [4], [5], [6], marking three orders-of-magnitude improvement over general-purpose processors [7] . These gains stem from specialized compute units, tightly coupled memory hierarchies, and optimized dataflows tailored to a narrow class of models. However, such accelerators typically rely on fixed memory access patterns, rigid compute pipelines, and bespoke data-movement hardware, which significantly limit their adaptability. As the deep neural network (DNN) landscape evolves rapidly—introducing new model architectures, operators, and optimization techniques—these fixed-function designs often suffer from suboptimal  \nVasilis Sakellariou and Hani Saleh are with the Computer and Information Engineering Department of Khalifa University, Abu Dhabi, UAE  \nVassilis Paliouras, Thanos Stouraitis, and Ioannis Kouretas are with the Electrical and Computer Engineering Department, University of Patras, Greece  \nutilization, particularly for workloads that deviate from their target assumptions [6] . Retargeting these accelerators to new applications requires extensive hardware redesign and verification effort, increasing development cost and deteriorating time-to-market [8] .  \nAn orthogonal approach to addressing flexibility has been the adoption of instruction-set-architecture(ISA)-based solutions, such as RISC-V [9], [10], [11], [12], [13], [14], [15] and commercial architectures with vector or matrix extensions (e.g., AMX, Armv8) . These systems improve programmability and design reuse by coupling general-purpose cores with tightly integrated accelerator engines (e.g., matrix multiplication and convolution units) via SIMD or vector extensions. While this approach enhances flexibility and provides a programming interface, it introduces area and power overheads associated with general-purpose execution and often struggles to match the performance of specializ","cbCaiob8fM9k9guI","https://ap.wps.com/l/cbCaiob8fM9k9guI","pdf",6892672,4,1,19,"English","en",105,"# Introduction\n## Motivation and limitations of fixed-function accelerators\n## ISA-based flexibility and representation tradeoffs\n## Proposed approach and contributions\n# Proposed Hardware Architecture","[{\"question\":\"What problem does the proposed ISA-based reconfigurable DNN accelerator aim to solve?\",\"answer\":\"It targets the mismatch between high-efficiency domain-specific accelerators and the flexibility needed for rapidly changing DNN architectures and operators, which fixed-function designs struggle to adapt to without redesign.\"},{\"question\":\"How does the proposed ISA improve flexibility without sacrificing efficiency?\",\"answer\":\"It enables fine-grained control over data movement, dynamic precision, and decoupled execution across data-fetching, tensor processing, and post-processing domains while maintaining high processing-element utilization with low control overhead.\"},{\"question\":\"What are the reported results for the Residue Number System instantiation?\",\"answer\":\"Using a 22-nm implementation with 3–8-bit dynamic precision, the architecture achieves 5.12–10.47 TOPS/W and up to 1.2× higher energy efficiency than a fixed-point counterpart, while preserving model accuracy and outperforming some state-of-the-art designs.\"}]","A Reconfigurable and Representation Adaptive ISA Based Architecture for Efficient DNN Acceleration | PDF",1784176948,48,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"a-reconfigurable-and-representation-adaptive-isa-based-architecture-for-efficient-dnn-acceleration","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/a-reconfigurable-and-representation-adaptive-isa-based-architecture-for-efficient-dnn-acceleration/81900/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the proposed ISA-based reconfigurable DNN accelerator aim to solve?","Question",{"text":76,"@type":77},"It targets the mismatch between high-efficiency domain-specific accelerators and the flexibility needed for rapidly changing DNN architectures and operators, which fixed-function designs struggle to adapt to without redesign.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed ISA improve flexibility without sacrificing efficiency?",{"text":81,"@type":77},"It enables fine-grained control over data movement, dynamic precision, and decoupled execution across data-fetching, tensor processing, and post-processing domains while maintaining high processing-element utilization with low control overhead.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the reported results for the Residue Number System instantiation?",{"text":85,"@type":77},"Using a 22-nm implementation with 3–8-bit dynamic precision, the architecture achieves 5.12–10.47 TOPS/W and up to 1.2× higher energy efficiency than a fixed-point counterpart, while preserving model accuracy and outperforming some state-of-the-art designs.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},"General","general"]