[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85999-en":3,"doc-seo-85999-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85999,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","From Self Attention to Connection Laplacian: A Unified Operator View of Transformers","Self-attention is treated as an operator-level geometric object for sequence models, where token features form a vector field over a token-position graph and attention operates as a connection walk. Messages are aggregated by a nonnegative walk matrix and transported along edges through learned linear maps. Single-head attention becomes an exact connection propagation step with constant transport, while multi-head attention becomes an edge-dependent walk with attention-gated headwise transports. The generator reduces to a random-walk connection Laplacian under specific stochastic, reversible, and metric-compatible conditions. Empirically, Transformers from 124M to 8B exhibit stable geometric operators in deeper layers and self-organized near scaled isometries, linking attention to classical geometric Laplacians.","arXiv :2607 . 10677v 1 [ cs .LG] 12 Jul 2026  \nFrom Self-Attention to Connection Laplacian  \nFrom Self-Attention to Connection Laplacian: A Unified Operator View of Transformers  \nBinbin Lin [binbinlin@zju.edu.cn](binbinlin@zju.edu.cn)  \nSchool of Software Technology, Zhejiang University, China  \nWei Chen [weichen.cw@zju.edu.cn](weichen.cw@zju.edu.cn)  \nCollege of Computer Science and Technology, Zhejiang University, China  \nYalun Li [yalunli@zju.edu.cn](yalunli@zju.edu.cn)  \nCollege of Computer Science and Technology, Zhejiang University, China  \nWenxiao Wang [wenxiaowang@zju.edu.cn](wenxiaowang@zju.edu.cn)  \nSchool of Software Technology, Zhejiang University, China  \nJieping Ye [yejieping.ye@alibaba-inc.com](yejieping.ye@alibaba-inc.com)  \nAlibaba Cloud, China  \nXiaofei He [xiaofeihe@cad.zju.edu.cn](xiaofeihe@cad.zju.edu.cn)  \nCollege of Computer Science and Technology, Zhejiang University, China  \nAbstract  \nSelf-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over the token-position graph and identify attention as a connection walk: messages are aggregated by a nonnegative walk matrix while being transported along each edge by a learned linear map. Within this framework, we prove that single-head attention (SHA) is exactly a connection propagation step with constant transport, and that multi-head attention (MHA) is exactly a single edge-dependent connection walk whose effective transport is an attentiongated mixture of headwise transports. We further clarify the conditions under which the corresponding generator reduces to a random-walk connection Laplacian, highlighting the roles of stochasticity, reversibility, and metric-compatible transports. Empirically, we find that trained Transformers across scales (from 124M to 8B) and structures (encoder/decoder) exhibit geometric structure consistent with our theory: effective attention graphs converge to stable geometric operators in deeper layers, learned transports self-organize into approximate scaled isometries, and both phenomena strengthen consistently with scale. Overall, the paper provides a precise connection-walk formalism that links self-attention to classical geometric operators, along with a set of operator-level tools for analyzing transformer models from a geometric perspective.  \nKeywords: connection Laplacian, self-attention, multi-head attention, transformers, geometric deep learning  \n1 Introduction  \nTransformers and their self-attention mechanism (Vaswani et al., 2017) have reshaped modern deep learning by enabling models to capture long-range dependencies across tokens. In self-attention, each token attends to other tokens through a learned similarity measure, producing a weighted aggregation of value vectors. The resulting attention matrix can be interpreted as a directed weighted graph where edges encode information flow. This paper  \nLin, Chen, Li, Wang, Ye and He  \ngives an operator-level geometric interpretation of self-attention: token representations forma vector field over the token-position graph (the token graph, for short), and attention acts as a connection walk that mixes tokens while transporting features along edges. Developing this operator view, we clarify the operator structure underlying multi-head attention (MHA), the depthwise behavior, and the geometry of token interactions.  \n1.1 Related work  \nExisting literature examines attention through various theoretical lenses, which we categorize into graph, tensor, dynamical, energy, kernel, and geometric perspectives.  \nGraph operator view. Self-attention can be viewed as a data-dependent linear operator that mixes token features via the attention matrix. This formulation aligns with graph transformers, where full attention implies all-to-all communication and sparse attention recovers local aggregation (Yun et al., 2019) . Analytic approaches often symmetrize or normalize","cbCairLZZsa5LveB","https://ap.wps.com/l/cbCairLZZsa5LveB","pdf",1815897,5,1,29,"English","en",105,"# Introduction\n## Related work","[{\"question\":\"How does the paper interpret self-attention geometrically at the operator level?\",\"answer\":\"It models tokens as a vector field over a token-position graph and interprets attention as a connection walk that aggregates messages and transports features along edges using learned linear maps.\"},{\"question\":\"What is the difference between the unified view of single-head attention and multi-head attention?\",\"answer\":\"Single-head attention is shown to be exactly a connection propagation step with constant transport, while multi-head attention corresponds to an edge-dependent connection walk whose transport is an attention-gated mixture of headwise transports.\"},{\"question\":\"When does the generator reduce to a random-walk connection Laplacian?\",\"answer\":\"It reduces under conditions involving stochasticity, reversibility, and metric-compatible transports, which the paper uses to identify the operator relationship to random-walk-style Laplacians.\"}]",1784207686,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"from-self-attention-to-connection-laplacian-a-unified-operator-view-of-transformers","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-self-attention-to-connection-laplacian-a-unified-operator-view-of-transformers/85999/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How does the paper interpret self-attention geometrically at the operator level?","Question",{"text":76,"@type":77},"It models tokens as a vector field over a token-position graph and interprets attention as a connection walk that aggregates messages and transports features along edges using learned linear maps.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the difference between the unified view of single-head attention and multi-head attention?",{"text":81,"@type":77},"Single-head attention is shown to be exactly a connection propagation step with constant transport, while multi-head attention corresponds to an edge-dependent connection walk whose transport is an attention-gated mixture of headwise transports.",{"name":83,"@type":74,"acceptedAnswer":84},"When does the generator reduce to a random-walk connection Laplacian?",{"text":85,"@type":77},"It reduces under conditions involving stochasticity, reversibility, and metric-compatible transports, which the paper uses to identify the operator relationship to random-walk-style Laplacians.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]