[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83645-en":3,"doc-seo-83645-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83645,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Elasticity in Parallel Sparse Triangular Solve","Elasticity in Parallel Sparse Triangular Solve presents stale-synchronous-parallel as an execution mode for solving sparse triangular linear systems and provides a general directed-acyclic-graph scheduler to generate such schedules. The approach enables overlap of synchronization and computation, yielding geometric-mean speedups of 7–30% for ElasticDivide over the synchronous GrowLocal scheduler on an ARM machine with 48 cores. On an x86 machine with 48 cores, geometric-mean speedups of 19–60% are reported over SpMP.","arXiv :2607 .02324v 1 [ cs .DC] 2 Jul 2026  \nELASTICITY IN PARALLEL SPARSE TRIANGULAR SOLVE  \nRAPHAEL S. STEINER, CHRISTOS K. MATZOROS, P´AL ANDR´AS PAPP, TONI B¨OHNLEIN,  \nAND ALBERT-JAN N. YZELMAN  \nAbstract. We introduce stale synchronous parallel as a mode of execution in parallel sparse triangular linear system solve and present a general directed-acyclic-graph scheduler capable of producing such schedules. Stale-synchronous-parallel schedules allow the overlap of synchronisation and compute which results in a geometric-mean speed-up of 7-30% of our scheduler, ElasticDivide, over state-of-the-art synchronous scheduler GrowLocal on an ARM machine using 48 cores. On an x86 machine using 48 cores, we report geometric-mean speedups of 19-60% over SpMP.  \n1. Introduction  \nSparse triangular solve is an omnipresent operation in computing. Whether it be in engineering, data analytics, artificial intelligence, or scientific computing, systems of equations need to be solved. This typically involves (sparse) triangular solve as part of the solving process directly, through methods like LU, QR, and Cholesky factorisations and Gauß–Seidel, or indirectly through (pre-)conditioning of the system for faster iterative solving methods such as conjugate gradient.  \nWhen solving ever bigger problem instances, sparsity plays an evermore important role. It allows one to cut redundant compute, but this comes at the cost of losing structure. This results in fine-grained dependencies and operations which makes sparse triangular solve (SpTrSV) a hard problem to parallelise. Therefore, getting the most performance out of modern multi-core architectures is a challenging endeavour.  \nTo combat this, SpTrSV performance is often optimised in larger contexts such as linear or symmetric solves where preprocessing allows for the reintroduction of structure. A prominent example of this is the nested dissection technique [Geo73, LRT79 , KK98 , APc04 , GBDD10 , BAvL+19] which concentrates non-zeroes for locality and introduces coarse-grained parallelism [DS05, LDS+23] . Another important technique is the introduction of so-called supernodes [CAGL+87 , Li05 , FZW+23 , LNP93 , YRE20 , SG04 , SGFS01] . Supernodes give rise to several optimisations. First, the grouping of operations into supernodes allows for the use of highly optimised dense kernels, and second, the dependency graph becomes coarser which enables dynamically computing a parallel schedule or dynamic dispatching methods as their benefits now outweigh their overhead. Another notable and generally applicable technique is the (recursive) splitting of SpTrSV into two SpTrSV problems of half the size and an easier-to-parallelise SpMV in-between [AS89, May09 , LLH+16 , LNL20 , AYU21] .  \nAfter applying these structural techniques (if they are feasible), one is left with a parallel scheduling problem. The tasks to be scheduled typically have irregular dependencies and their size can range from whole blocks or supernodes to a couple of floating-point operations. It is this scheduling problem that we address in this paper. More precisely, we consider the fine-grained scheduling problem of parallelising the forward-substitution algorithm, cf. Algorithm 2.1. Nevertheless, our algorithm and the involved ideas may, of course, be applied to any parallel scheduling problem.  \nTo get the most performance out of SpTrSV, a parallel scheduling algorithm must  \n• balance workloads among cores,  \n• limit coordination overhead, and  \n• consider spatial and temporal locality.  \nKey words and phrases. Sparse triangular linear system solve, SpTrSV, SpTrSM, forward-and backwardsubstitution algorithm, stale-synchronous-parallel algorithm.  \nELASTICITY IN PARALLEL SPARSE TRIANGULAR SOLVE 2  \n HDagg  GrowLocal  ElasticDivide  \nFigure 1 .1. Geometric-mean speed-ups over Serial for various data sets, cf. §5.2, on ARM Kunpeng with 10-th to 90-th percentile range shown.  \nSeveral algorithms to this end have been proposed in the ","cbCaia6CaszxFF4g","https://ap.wps.com/l/cbCaia6CaszxFF4g","pdf",1130745,5,1,23,"English","en",105,"# Introduction\n## Our contribution","[{\"question\":\"What execution model and scheduling approach does the paper introduce for sparse triangular solve?\",\"answer\":\"The paper uses stale-synchronous-parallel execution and introduces a general directed-acyclic-graph (DAG) scheduler capable of producing such schedules.\"},{\"question\":\"How does stale-synchronous parallelism improve performance in sparse triangular solve?\",\"answer\":\"It overlaps synchronization with computation, leading to measurable geometric-mean speedups versus established synchronous and asynchronous schedulers.\"},{\"question\":\"What performance improvements are reported on ARM and x86 systems?\",\"answer\":\"On an ARM machine with 48 cores, ElasticDivide achieves geometric-mean speedups of 7–30% over GrowLocal, while on x86 with 48 cores it reports 19–60% speedups over SpMP.\"}]",1784189474,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"elasticity-in-parallel-sparse-triangular-solve","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/elasticity-in-parallel-sparse-triangular-solve/83645/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What execution model and scheduling approach does the paper introduce for sparse triangular solve?","Question",{"text":76,"@type":77},"The paper uses stale-synchronous-parallel execution and introduces a general directed-acyclic-graph (DAG) scheduler capable of producing such schedules.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does stale-synchronous parallelism improve performance in sparse triangular solve?",{"text":81,"@type":77},"It overlaps synchronization with computation, leading to measurable geometric-mean speedups versus established synchronous and asynchronous schedulers.",{"name":83,"@type":74,"acceptedAnswer":84},"What performance improvements are reported on ARM and x86 systems?",{"text":85,"@type":77},"On an ARM machine with 48 cores, ElasticDivide achieves geometric-mean speedups of 7–30% over GrowLocal, while on x86 with 48 cores it reports 19–60% speedups over SpMP.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]