[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83648-en":3,"doc-seo-83648-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83648,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Coalesced Matrix-Free Finite Elements in Cell-Wise Storage","GPU-oriented high-order continuous finite elements are formulated so the redundant cell-wise (element-local) vector remains the persistent primary representation of field data, eliminating the need to treat it as a transient stage in matrix-free operator evaluation. With an appropriate continuous-image preconditioner, flexible conjugate gradient iterations run exactly on the unassembled representation, matching Krylov scalars of assembled solves while restricting inter-element communication to the preconditioner. Direct stiffness summation is implemented via an axis-split sequence of one-to-one face exchanges, avoiding gatherscatter, atomics, and coloring. Cell-wise storage also enables benchmarking blocked memory layouts optimized for GPU throughput, and experiments show superior solver performance over state-of-the-art matrix-free methods.","arXiv :2607 .02335v1 [math .NA] 2 Jul 2026  \nCoalesced Matrix-Free Finite Elements in Cell-Wise Storage  \nMichal Wichrowski  \nAbstract  \nWe present a GPU-oriented formulation of continuous high-order finite elements in which the redundant, cell-wise (element-local) vector is the persistent primary representation of all field data, rather than a transient stage of matrix-free operator evaluation. We prove that, given apreconditioner whose image is continuous, the entire flexible conjugate gradient iteration can be carried out exactly on this unassembled representation: a simple primal-dual pairing identity shows that all Krylov scalars computed from local data coincide with those of the assembled solve, so inter-element communication is confined entirely to the preconditioner. The required direct stiffness summation (DSS) is then realized without indirect gather-scatter, atomics, or coloring, by a dimensionally-split cascade of one-to-one face exchanges that provably accumulates edge and vertex contributions as a byproduct of sequential axis passes; unstructured macro-block interfaces and h-adaptive hanging nodes are handled by disjoint topological kernelsand a shadow-cell wrapper that leaves the high-throughput sweeps untouched. The cell-wise storage decouples the memory layout from the mesh topology, and we exploit this freedom to benchmark blocked layouts that trade memory coalescing against element contiguity. Numerical experiments on modern GPUs demonstrate that the resulting operator evaluation and solver outperform state-of-the-art matrix-free implementations, signifficantly exceeding throughput of existing implementations.  \nKeywords: Direct stiffness summation, matrix-free, finite elements, multigrid, GPU  \nAMS subject classifications: 65Y10, 65Y20, 65N55, 65N30  \n1 Introduction  \nMatrix-free evaluation of high-order finite element operators has become the method of choice for large-scale simulation on modern hardware. Instead of assembling and storing a sparse matrix, the operator action is recomputed on the fly from element-local data using sum factorization, reducing both memory footprint and memory traffic by an order of magnitude at moderate polynomial degrees [28, 11, 21, 5] . This shift is driven by the hardware itself: over the past two decades the floating-point throughput of processors has grown far faster than their memory bandwidth, so that virtually all sparse and element-wise operations are limited by data movement rather than arithmetic [34] . The trend is most pronounced on GPUs, which now dominate the upper end of the TOP500 list and deliver their nominal bandwidth only under stringent conditions on the access pattern.  \nGPUs add two further twists. First, their memory subsystem is throughput-oriented: peak bandwidth is attained only when the threads of a warp access contiguous, aligned memory (coalescing), and any indirection-based access pattern forfeits a substantial fraction of it. Second, anever-growing share of the silicon is devoted to matrix units (tensor cores) and to reduced-precision  \n0 Interdisziplin¨ares Zentrum f¨ur Wissenschaftliches Rechnen (IWR), Ruprecht-Karls-Universit¨at Heidelberg, Germany, [mwichro@mimuw.edu.pl](mwichro@mimuw.edu.pl)  \narithmetic, to the point that double-precision capability stagnates or is emulated through lowerprecision units [24] . An algorithm aiming to maximize hardware utilization should therefore (i) arrange its data so that every load and store is coalesced, (ii) cast its inner kernels as small dense matrix products, and (iii) tolerate reduced or mixed precision inside the preconditioner [19] . These three requirements shape the design presented in this paper.  \nState-of-the-art matrix-free frameworks: deal.II [21, 22, 3], MFEM [2], libCEED [8, 1], Nek5000/NekRS [14,  \n27], and atmospheric models like NUMA [25] organize the operator action around an assembled global vector of unique degrees of freedom. Every operator application then begins by gathering ","cbCaifC1c9QYkxRz","https://ap.wps.com/l/cbCaifC1c9QYkxRz","pdf",606442,4,1,31,"English","en",105,"# Introduction\n## GPU-oriented matrix-free finite elements\n## Motivation: bandwidth and coalesced memory access\n## Alternative solver perspective: cell-wise persistent representation\n## Background and distinguishing contribution","[{\"question\":\"What is the main idea behind the GPU-oriented formulation in this work?\",\"answer\":\"The method keeps a redundant cell-wise (element-local) vector as the persistent primary representation of field data, so it is used directly by the solver rather than as a transient stage in matrix-free operator evaluation.\"},{\"question\":\"How does the approach ensure Krylov scalars match those of an assembled solve?\",\"answer\":\"With a preconditioner whose image is continuous, a primal-dual pairing identity shows that Krylov scalars computed from local cell-wise data coincide exactly with those from the assembled solve; communication between elements is confined to the preconditioner.\"},{\"question\":\"How is direct stiffness summation (DSS) performed without traditional GPU bottlenecks?\",\"answer\":\"DSS is realized without gather-scatter, atomics, or coloring by using a dimensionally split cascade of one-to-one face exchanges that accumulates edge and vertex contributions during sequential axis passes, with additional handling for unstructured macro-block interfaces and h-adaptive hanging nodes.\"}]",1784189502,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"coalesced-matrix-free-finite-elements-in-cell-wise-storage","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/coalesced-matrix-free-finite-elements-in-cell-wise-storage/83648/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main idea behind the GPU-oriented formulation in this work?","Question",{"text":75,"@type":76},"The method keeps a redundant cell-wise (element-local) vector as the persistent primary representation of field data, so it is used directly by the solver rather than as a transient stage in matrix-free operator evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the approach ensure Krylov scalars match those of an assembled solve?",{"text":80,"@type":76},"With a preconditioner whose image is continuous, a primal-dual pairing identity shows that Krylov scalars computed from local cell-wise data coincide exactly with those from the assembled solve; communication between elements is confined to the preconditioner.",{"name":82,"@type":73,"acceptedAnswer":83},"How is direct stiffness summation (DSS) performed without traditional GPU bottlenecks?",{"text":84,"@type":76},"DSS is realized without gather-scatter, atomics, or coloring by using a dimensionally split cascade of one-to-one face exchanges that accumulates edge and vertex contributions during sequential axis passes, with additional handling for unstructured macro-block interfaces and h-adaptive hanging nodes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]