[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134468-en":3,"doc-seo-134468-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},134468,5909887254083,"\tWilliam","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","FRED - A Wafer-scale Fabric for 3D Parallel DNN Training","Wafer-scale systems tightly integrate high-end accelerator chiplets with high-speed wafer-scale interconnects, offering low-latency, high-bandwidth connectivity for deep neural network training. Existing network-on-wafer designs such as 2D Meshes lack the flexibility to match diverse parallelization patterns. This work introduces Fred, a wafer-scale fabric using tiny microswitches to form a distributed nonblocking on-wafer topology with in-switch collective support for arbitrary accelerator groups. Benchmarks show sizable end-to-end training-time reductions for ResNet-152, Transformer-17B, GPT-3, and Transformer-1T versus a wafer-scale Mesh baseline.","FRED: A Wafer-scale Fabric for 3D Parallel DNN Training  \nSaeed Rashidi∗ [saeed.rashidi@gatech.edu](saeed.rashidi@gatech.edu)[ ](saeed.rashidi@gatech.edu)Georgia Institute of Technology Atlanta, Georgia, USA  \nWilliam Won  \n[william.won@gatech.edu](william.won@gatech.edu)[ ](william.won@gatech.edu)Georgia Institute of Technology Atlanta, Georgia, USA  \nSudarshan Srinivasan  \n[sudarshan.srinivasan@intel.com](sudarshan.srinivasan@intel.com)[ ](sudarshan.srinivasan@intel.com)Intel Bangalore, Karnataka, India  \nPuneet Gupta  \n[puneetg@ucla.edu](puneetg@ucla.edu)[ ](puneetg@ucla.edu)UCLA  \nLos Angeles, California, USA  \nTushar Krishna  \n[tushar@ece.gatech.edu](tushar@ece.gatech.edu)[ ](tushar@ece.gatech.edu)Georgia Institute of Technology Atlanta, Georgia, USA  \narXiv :2406 . 19580v2 [ cs .AR] 9 Jun 2025  \nAbstract  \nWafer-scale systems are an emerging technology that tightly integrates high-end accelerator chiplets with high-speed wafer-scale interconnects, enabling low-latency and high-bandwidth connectivity. This makes them a promising platform for deep neural network (DNN) training. However, current network-on-wafer topologies, such as 2D Meshes, lack the flexibility needed to support various parallelization strategies effectively. In this paper, we propose Fred, a wafer-scale fabric architecture tailored to the unique communication needs of DNN training. Fred creates a distributed on-wafer topology with tiny microswitches, providing nonblocking connectivity for collective communications between arbitrary groups of accelerators and enabling in-switch collective support. Our results show that for sample parallelization strategies, Fred can improve the average end-to-end training time of ResNet-152, Transformer-17B, GPT-3, and Transformer-1T by 1.76×, 1.87×, 1.34×, and 1.4×, respectively, compared to a baseline wafer-scale Mesh.  \nCCS Concepts  \n• Hardware → Emerging technologies; • Networks → Network architectures; • Computer systems organization → Architectures.  \nKeywords  \ndistributed training, wafer-scale platforms  \nACM Reference Format:  \nSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta, and Tushar Krishna. 2025. FRED: A Wafer-scale Fabric for 3D Parallel DNN Training. In Proceedings of The 52nd IEEE/ACM International Symposium on Computer Architecture (’ISCA 2025’) . ACM, New York, NY, USA, 15 pages. [https://doi.org/3695053.3731055](https://doi.org/3695053.3731055)  \n1 Introduction  \nDNN models are on an exponential growth curve. A recent study shows that in less than two years, the compute and memory requirements for DNN training have increased by 1,800× and 1,500×, respectively [58] . Distributing the training across multiple accelerators or neural processing units (NPUs) is a common practice today to reduce the training time. However, one critical side effect  \n∗The author is now at Meta: [rashidi1saeed@meta.com](rashidi1saeed@meta.com)  \n’ISCA 2025’, June 21–25, 2025, Tokyo, Japan 2025. ACM ISBN 978-1-4503-XXXX-X/18/06 [https://doi.org/3695053.3731055](https://doi.org/3695053.3731055)  \nof distributed training is the communication overhead between NPUs to synchronize model gradients and/or activations, depending on the parallelization strategy. As the number of NPUs scales, communication overhead increases, up to a point where it becomes the dominant factor in distributed training latency [26, 27, 54, 55] .  \nThere are fundamental limits to the bandwidth that can be provided even by high-speed rack-scale fabrics (such as NVLink [33]), and thus there has been a growing interest in platforms that integrate multiple NPUs together in the same package. Cerebras [10] demonstrated one extreme incarnation of this idea in the form of a monolithic wafer with NPUs connected to one another. More costand yield-effective approaches include silicon/organic interposerbased approaches [24, 59] or using Silicon Interconnect Fabric (SiIF), which bonds chiplets directly onto a full thickness silicon wafer without needing a","cbCaihWAjCgYpAYB","https://ap.wps.com/l/cbCaihWAjCgYpAYB","pdf",1289421,5,1,15,"English","en",105,"# Introduction\n## DNN scaling and distributed training overhead\n## Wafer-scale substrate and fabric topologies\n## Parallelism strategies for DNN training","[{\"question\":\"Why does communication become a bottleneck in distributed DNN training?\",\"answer\":\"As the number of NPUs increases, the communication overhead required to synchronize gradients and/or activations grows, eventually becoming the dominant contributor to distributed training latency.\"},{\"question\":\"What limitation of 2D Mesh fabrics motivates the Fred architecture?\",\"answer\":\"2D Mesh topologies are inherently blocking, making them inefficient for DNN training communication patterns that require flexible, strategy-dependent connectivity.\"},{\"question\":\"How does Fred support collective communication during training?\",\"answer\":\"Fred builds a distributed on-wafer topology with tiny microswitches to provide nonblocking connectivity and enabling in-switch collective support for arbitrary accelerator groups.\"}]","FRED - A Wafer-scale Fabric for 3D Parallel DNN Training | PDF",1787263274,38,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"fred-a-wafer-scale-fabric-for-3d-parallel-dnn-training","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/fred-a-wafer-scale-fabric-for-3d-parallel-dnn-training/134468/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-25","2026-08-20",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why does communication become a bottleneck in distributed DNN training?","Question",{"text":77,"@type":78},"As the number of NPUs increases, the communication overhead required to synchronize gradients and/or activations grows, eventually becoming the dominant contributor to distributed training latency.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What limitation of 2D Mesh fabrics motivates the Fred architecture?",{"text":82,"@type":78},"2D Mesh topologies are inherently blocking, making them inefficient for DNN training communication patterns that require flexible, strategy-dependent connectivity.",{"name":84,"@type":75,"acceptedAnswer":85},"How does Fred support collective communication during training?",{"text":86,"@type":78},"Fred builds a distributed on-wafer topology with tiny microswitches to provide nonblocking connectivity and enabling in-switch collective support for arbitrary accelerator groups.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]