[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117571-en":3,"doc-seo-117571-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117571,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","The underlying structures of self-attention - symmetry, directionality, and emergent dynamics in Transformer training","Self-attention is fundamental to Transformer architectures, yet the way information is encoded in self-attention matrices and how training objectives shape this encoding remains insufficiently understood. The work introduces a mathematical framework that derives the structures behind self-attention weight updates. Bidirectional training yields symmetry in weight matrices, while autoregressive training produces directionality and column dominance. Results hold across multiple Transformer models and modalities, and symmetric initialization improves encoder-only performance on language tasks.","The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training  \nMatteo Saponati 1 * Pascal Sager 2 1 * Pau Vilimelis Aceituno 1 Thilo Stadelmann 2 3 Benjamin Grewe 1 4  \nAbstract  \nSelf-attention is essential to Transformer architectures, yet how information is embedded in the self-attention matrices and how different objective functions impact this process remains unclear.  \nWe present a mathematical framework to analyze self-attention matrices by deriving the structures governing their weight updates. Using this framework, we demonstrate that bidirectional training induces symmetry in the weight matrices, while autoregressive training results in directionality and column dominance. Our theoretical findings are validated across multiple Transformer models  \n—including ModernBERT, GPT, LLaMA3, and Mistral—and input modalities like text, vision, and audio. Finally, we apply these insights by showing that symmetric initialization improves the performance of encoder-only models on language tasks. This mathematical analysis offers a novel theoretical perspective on how information is embedded through self-attention, thereby improving the interpretability of Transformer models.  \n1. Introduction  \nTransformer models now achieve state-of-the-art performance across a wide range of tasks and domains (Radford et al., 2019 ; Dosovitskiy et al., 2021 ; Radford et al., 2023) . Despite their success, the internal mechanisms governing their decision-making processes remain poorly understood, raising concerns regarding model alignment, reliability, and safety (Wang et al., 2023 ; Yao et al., 2024) . A key challenge  \n*Equal contribution 1Institute of Neuroinformatics, ETH Z¨urichand University of Z¨urich, Z¨urich, Switzerland 2 Centre for Artificial Intelligence, Z¨urich University of Applied Sciences, Winterthur, Switzerland 3ECLT European Centre for Living Technology, Venice, Italy 4ETH AI Center, Z¨urich, Switzerland. Correspondence to: Matteo Saponati \u003C[masapo@ini.ethz.ch](masapo@ini.ethz.ch) >, Benjamin Grewe \u003C[bgrewe@ini.ethz.ch](bgrewe@ini.ethz.ch) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nin understanding these models is unraveling the structures of self-attention, which is essential to Transformer architectures. Current literature largely overlooks the nature of the self-attention weight matrices during autoregressive training, where the model predicts the next token in a sequence given previous ones (Radford et al., 2019 ; Black et al., 2021 ; Touvron et al., 2023) and bidirectional training, where the model predicts a missing token given the full sequence (Devlin et al., 2019 ; Bao et al., 2022 ; Warner et al., 2024) . Understanding self-attention requires answering two fundamental questions: How can we interpret the structures learned in the self-attention matrices? What is the impact of different objective functions on these matrices?  \nPrevious work used sparse auto-encoders to identify interpretable features (Huben et al., 2024 ; Bricken et al., 2023), circuit analysis to interpret Transformer components (Olah et al., 2020 ; Elhage et al., 2021 ; Olah, 2022), and techniques like the logit lens to analyze self-attention mechanisms (Geva et al., 2021 ; Dar et al., 2023) (for a detailed discussion, see Section 5) . However, these methods do not reveal the structural patterns in self-attention matrices or the transformations they encode. Crucially, how autoregressive and bidirectional training shape specific weight structures remains unclear.  \nTo address this gap, we introduce a novel framework for analyzing self-attention matrices and understanding how different objective functions define their weight updates. We then use this framework to derive understandable mathematical structures that should emerge from such updates. Finally, we verify these interpretable structures","cbCairZmr7z9lK5a","https://ap.wps.com/l/cbCairZmr7z9lK5a","pdf",1112348,1,37,"English","en",105,"# Introduction\n## Autoregressive vs bidirectional training\n## Related work and analysis gaps\n# Autoregressive and bidirectional training leads to directional and symmetric weight updates\n## Interpreting self-attention with bilinear forms","[{\"question\":\"What problem does the paper address about self-attention matrices?\",\"answer\":\"It studies how information is embedded in self-attention matrices and how different objective functions affect the structures learned during Transformer training.\"},{\"question\":\"How do bidirectional and autoregressive training differ in the paper’s framework?\",\"answer\":\"Bidirectional training induces symmetry in the self-attention weight matrices, while autoregressive training produces directionality and column dominance.\"},{\"question\":\"How are the theoretical findings validated and applied?\",\"answer\":\"They are tested numerically across many Transformer models and modalities, and symmetric initialization is shown to improve performance of encoder-only models on language tasks.\"}]","The underlying structures of self-attention - symmetry, directionality, and emergent dynamics in Transformer training | PDF",1785677050,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-underlying-structures-of-self-attention-symmetry-directionality-and-emergent-dynamics-in-transformer-training","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-underlying-structures-of-self-attention-symmetry-directionality-and-emergent-dynamics-in-transformer-training/117571/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address about self-attention matrices?","Question",{"text":75,"@type":76},"It studies how information is embedded in self-attention matrices and how different objective functions affect the structures learned during Transformer training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do bidirectional and autoregressive training differ in the paper’s framework?",{"text":80,"@type":76},"Bidirectional training induces symmetry in the self-attention weight matrices, while autoregressive training produces directionality and column dominance.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the theoretical findings validated and applied?",{"text":84,"@type":76},"They are tested numerically across many Transformer models and modalities, and symmetric initialization is shown to improve performance of encoder-only models on language tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]