[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84689-en":3,"doc-seo-84689-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84689,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage","Geometric multigrid preconditioning is developed for high-order continuous finite elements using only redundant, cell-wise stored vectors, avoiding formation of the assembled global vector on any hierarchy level. In this setting, hanging-node constraints are never explicitly assembled: tensor-product transfer operators applied to the unassembled residual reproduce classical constrained restriction including the transposed constraint-matrix action, and local smoothing edge operators reduce to pointwise residual masking. The resulting cell-wise V-cycle is proved equivalent to classical local multigrid, retaining its convergence theory. Numerical tests for the Laplace operator confirm grid-independent convergence and high GPU throughput.","arXiv :2607 .03413v1 [math .NA] 3 Jul 2026  \nCoalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise  \nStorage  \nMichal Wichrowski  \nAbstract  \nWe present a geometric multigrid preconditioner for high-order continuous finite elements that operates entirely on redundant, cell-wise stored vectors: the assembled global vector is never formed on any level of the hierarchy. In this storage paradigm the machinery that classically complicates adaptive multigrid dissolves. Hanging-node constraints are never assembled:  \nwe prove that the plain tensor-product transfer operators, applied to the unassembled residual, algebraically reproduce the classical constrained restriction, including the action of the transposed constraint matrix, and the edge operators of local smoothing reduce to a pointwise masking of the residual, with no splitting of the level operator into interior and edge blocks.  \nAs a consequence, the single inter-cell primitive of the whole V-cycle can use a topologically oblivious structured kernel even on adaptively refined meshes. We prove that the resulting cell-wise V-cycle is equivalent, iterate by iterate, to the classical local multigrid method, and therefore inherits its convergence theory. Numerical experiments for the Laplace operator confirm grid-independent convergence that is essentially unaffected by local refinement; on a single GPU, using nothing more than a masked point-Jacobi smoother, the solver sustains up to 1 .1 GDoF/s per V-cycle in double precision and reaches end-to-end solve throughput on par with patch-smoother-based solvers.  \nKeywords: geometric multigrid, adaptive mesh refinement, hanging nodes, matrix-free, finite elements, GPU  \nAMS subject classifications: 65N55, 65N30, 65F08, 65Y10, 65Y20  \n1 Introduction  \nGeometric multigrid is the natural solver for high-order finite element discretizations of elliptic problems: with sum-factorized operator evaluation and tensor-product inter-grid transfers, the cost of a V-cycle scales as O (pd+1) per cell and level-independent convergence is well established for a broad class of problems [17, 4, 30, 33] . On modern GPUs, where virtually all element-wise operations are limited by data movement rather than arithmetic [41], the matrix-free realization of such solvers—including the block-structured hierarchical-hybrid-grid solvers [3, 15] and the hybrid geometric/polynomial hierarchies used at scale [13, 22]—has been refined to the point that the sum-factorized kernels themselves are no longer the bottleneck: the cost is concentrated in how data is stored and moved between them [14, 23] .  \nAdaptive mesh refinement compounds this problem. On locally refined quadrilateral and hexahedral meshes, conformity is classically enforced by constraining the hanging nodes that appear along boundaries between regions of different refinement depth to interpolate their conforming  \n0 Interdisziplin¨ares Zentrum f¨ur Wissenschaftliches Rechnen (IWR), Ruprecht-Karls-Universit¨at Heidelberg, Germany, [mwichro@mimuw.edu.pl](mwichro@mimuw.edu.pl)  \nneighbors, and these constraints reach into every component of the solver. In an assembled framework they are built into the assembly loop, eliminate rows and columns of the system matrix, and are applied explicitly during every operator evaluation, every smoother step, and every inter-grid transfer [1, 20] . Modern matrix-free solvers avoid forming the stiffness matrix but retain the full constraint infrastructure: index sets, constraint matrices, and the splitting of level operators into interior and edge blocks required by the local smoothing algorithm [18, 21, 27, 24] . On a GPU this infrastructure is not merely an implementation nuisance. Constraint resolution is realized through indirect, data-dependent memory accesses embedded in the gather-scatter stage—exactly the access pattern that forfeits coalescing and fragments the high-throughput sweeps that the structured parts of the mesh would otherwise admit. The h","cbCaiqmGKrKjRGhC","https://ap.wps.com/l/cbCaiqmGKrKjRGhC","pdf",621788,1,33,"English","en",105,"# Introduction\n## Motivation: Matrix-free multigrid on GPUs\n## Challenge: Adaptive refinement and hanging-node constraints\n## Prior theory and implementations\n## Core contribution: cell-wise storage eliminating constraint machinery\n## Paper outline (implied)","[{\"question\":\"What key idea enables the proposed multigrid to avoid assembling global vectors and constraints?\",\"answer\":\"The solver uses a persistent cell-wise (redundant) storage framework so the assembled global vector is never formed, and hanging-node constraints are not explicitly assembled as algebraic objects.\"},{\"question\":\"How does the method handle hanging-node constraints without assembling them?\",\"answer\":\"Plain tensor-product transfer operators applied to the unassembled residual algebraically reproduce classical constrained restriction, including the effect of the transposed constraint matrix, while edge smoothing reduces to pointwise masking of the residual.\"},{\"question\":\"What performance and convergence properties are demonstrated on GPUs?\",\"answer\":\"Experiments for the Laplace operator show grid-independent convergence essentially unaffected by local refinement, and on a single GPU a masked point-Jacobi smoother sustains up to about 1.1 GDoF/s per V-cycle in double precision with end-to-end throughput comparable to patch-smoother solvers.\"}]",1784197673,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"coalesced-matrix-free-geometric-multigrid-on-persistent-cell-wise-storage","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/coalesced-matrix-free-geometric-multigrid-on-persistent-cell-wise-storage/84689/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key idea enables the proposed multigrid to avoid assembling global vectors and constraints?","Question",{"text":75,"@type":76},"The solver uses a persistent cell-wise (redundant) storage framework so the assembled global vector is never formed, and hanging-node constraints are not explicitly assembled as algebraic objects.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method handle hanging-node constraints without assembling them?",{"text":80,"@type":76},"Plain tensor-product transfer operators applied to the unassembled residual algebraically reproduce classical constrained restriction, including the effect of the transposed constraint matrix, while edge smoothing reduces to pointwise masking of the residual.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance and convergence properties are demonstrated on GPUs?",{"text":84,"@type":76},"Experiments for the Laplace operator show grid-independent convergence essentially unaffected by local refinement, and on a single GPU a masked point-Jacobi smoother sustains up to about 1.1 GDoF/s per V-cycle in double precision with end-to-end throughput comparable to patch-smoother solvers.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]