[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81596-en":3,"doc-seo-81596-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81596,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Trivance Latency-Optimal AllReduce by Shortcutting Multiport Networks","AllReduce is a core collective communication primitive whose latency is driven by the number of communication steps and by the effective communication distance, especially in latency-sensitive regimes. Topologies such as direct-connect tori can suffer large distances due to limited bisection bandwidth. Trivance introduces a latency-optimal AllReduce completing within log3 steps, improving 50% over Swing and Recursive Doubling, cutting congestion by 3× versus Bruck while preserving bandwidth optimality. It uses joint reductions with both ports of a bidirectional ring per step, and generalizes naturally to multidimensional torus networks.","Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks  \nAnton Juerss  \nWeizenbaum Institute & TU Berlin  \nVamsi Addanki Purdue University  \nStefan Schmid  \nTU Berlin & Weizenbaum Institute  \narXiv :2602 . 17254v2 [ cs .DC] 9 Jul 2026  \nAbstract  \nAllReduce is a fundamental collective communication operation in distributed computing and a key performance bottleneck for large-scale training and inference. Its completion time is determined by the number of communication steps, which dominate latency-sensitive workloads, and the communication distance affecting both latency-and bandwidthbound regimes. Direct-connect topologies, such as Google’s TPUv4 tori, are particularly prone to large communication distances due to limited bisection bandwidth.  \nIn this paper, we present Trivance, a novel AllReduce algorithm that completes within log3 􀀽 steps—a 50% improvement in comparison to Swing and Recursive Doubling, while reducing congestion compared to Bruck’s algorithm by a factor of three and preserving bandwidth-optimality. Trivance exploits both transmission ports ofa bidirectional ring within each step to triple the communication distance along both directions simultaneously. By performing joint reductions, Trivance improves both the number of steps and network congestion. We further show that Trivance extends naturally to multidimensional torus networks, retaining its latency advantage while achieving performance comparable to bandwidth-optimal algorithms for large AllReduce sizes.  \nOur packet-level SST simulation shows that Trivance improves state-of-the-art approaches by 5-30% for AllReduce sizes up to 8 MiB, in high-bandwidth settings up to 32 MiBand for 3D tori up to 128 MiB. Throughout the evaluation, Trivance remains the best-performing latency-optimal algorithm.  \n1 Introduction  \nCollective communication lies at the heart of many highperformance computing applications and machine learning (ML), both for training and inference. As ML model sizes continue to grow [3, 23, 31], efficient collective communication primitives are critical for minimizing training time and maximizing hardware utilization [5, 23] . Recent research has focused on improving collective communication performance in real networks, spanning general algorithm design [10, 28, 29], synthesizing custom algorithms for specific topologies [5, 22, 30], and building scalable network topologies for large-scale training [19, 36] .  \nAllReduce stands out as a fundamental collective primitive, aggregating and disseminating data such as gradients  \nacross nodes during parallel training [9, 18, 29]. AllReduce often dominates in both HPC and deep learning environments where the latter heavily relies on AllReduce for gradient synchronization. In production HPC workloads, measurementson large-scale systems have shown that more than 40% of the total Message Passing Interface (MPI) time can be spent in MPI_AllReduce and MPI_Reduce combined [8, 25, 34] .  \nWith the adoption of torus topologies in large-scale GPU clusters [19, 36], optimizing AllReduce for multiport, bidirectional networks is key to minimizing communication overheads [29] . In such networks, each node connects to its immediate neighbors via bidirectional links, enabling it to send and receive messages concurrently. The structural simplicity and cost-efficiency of torus topologies make them a compelling choice for large-scale deployments, particularly in the context of machine learning workloads [14] . For this reason, several accelerator systems adopt torus-like topologies, most notably Google’s TPU platforms: TPUv4 organizes chips into 4×4×4 cubes using fixed electrical ICI links and connects up to 64 such cubes through same bandwidth optical ICI links routed by a reconfigurable circuit-switched fabric, forming 3D tori of up to 4,096 chips in different shapes [19] . More recently, Google’s TPU 8t scales a single superpod to as many as 9,600 chips [12] .  \nCollective communication typically ","cbCaifOTLk2ZLzhV","https://ap.wps.com/l/cbCaifOTLk2ZLzhV","pdf",1138363,5,1,20,"English","en",105,"# Introduction\n## Background on collective communication and AllReduce\n## Network topology and bidirectional multiport settings\n## Communication steps, latency, and congestion trade-offs\n## Related bounds and existing algorithms\n## Motivation via examples and congestion differences","[{\"question\":\"Why does AllReduce latency depend on communication steps and distance?\",\"answer\":\"AllReduce completion time is determined by the number of communication steps, which dominate latency-sensitive workloads, and by communication distance that affects both latency- and bandwidth-bound regimes.\"},{\"question\":\"What performance improvements does Trivance claim over existing AllReduce algorithms?\",\"answer\":\"Trivance completes within log3 steps, a 50% improvement over Swing and Recursive Doubling, reduces congestion by a factor of three compared to Bruck, and keeps bandwidth optimality.\"},{\"question\":\"How does Trivance exploit multiport networks to reduce steps and congestion?\",\"answer\":\"Within each step, Trivance uses both transmission ports of a bidirectional ring to triple communication distance in both directions simultaneously, and performs joint reductions to improve both step count and network congestion.\"}]",1784174601,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"trivance-latency-optimal-allreduce-by-shortcutting-multiport-networks","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/trivance-latency-optimal-allreduce-by-shortcutting-multiport-networks/81596/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does AllReduce latency depend on communication steps and distance?","Question",{"text":76,"@type":77},"AllReduce completion time is determined by the number of communication steps, which dominate latency-sensitive workloads, and by communication distance that affects both latency- and bandwidth-bound regimes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What performance improvements does Trivance claim over existing AllReduce algorithms?",{"text":81,"@type":77},"Trivance completes within log3 steps, a 50% improvement over Swing and Recursive Doubling, reduces congestion by a factor of three compared to Bruck, and keeps bandwidth optimality.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Trivance exploit multiport networks to reduce steps and congestion?",{"text":85,"@type":77},"Within each step, Trivance uses both transmission ports of a bidirectional ring to triple communication distance in both directions simultaneously, and performs joint reductions to improve both step count and network congestion.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]