[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123525-en":3,"doc-seo-123525-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123525,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Counting in Small Transformers - The Delicate Interplay between Attention and Feed-Forward Layers - Abstract","Architectural choices in transformer blocks strongly affect the solution space beyond scaling. This work analyzes how simple transformer blocks solve a histogram task: counting occurrences of each token in input sequences. Despite its apparent simplicity, the task exposes interactions among predictive performance, vocabulary and embedding sizes, token-mixing mechanisms, and feed-forward layer capacity. Two strategies emerge—relation-based and inventory-based counting—allocating functionality differently across attention and feed-forward layers, with robustness improvements from softmax and beginning-of-sequence tokens, confirmed by mechanistic inspection during training.","Counting in Small Transformers:  \nThe Delicate Interplay between Attention and Feed-Forward Layers  \nFreya Behrens 1 Luca Biggio 2 Lenka Zdeborov 1  \nAbstract  \nNext to scaling considerations, architectural design choices profoundly shape the solution space of transformers. In this work, we analyze the solutions simple transformer blocks implement when tackling the histogram task: counting items in sequences. Despite its simplicity, this task reveals a complex interplay between predictive performance, vocabulary and embedding sizes, tokenmixing mechanisms, and feed-forward layer capacity. We identify two theoretical counting strategies transformers adopt, relation-based and inventory-based counting, each defining distinct learning regimes for the task. These strategies dictate how functionality is distributed between attention and feed-forward layers. We further show that adding softmax and beginning-of-sequence tokens allow for more robustness when embedding dimensions are comparatively small. Empirical introspection of trained models closely confirms both the learning regimes of the various architectures and the formation of these strategies during training. We demonstrate how a basic task that requires only aggregation and selection is significantly impacted by minor design changes.  \n1. Introduction  \nTransformers are the key neural network behind many recent deep learning advances, most notably large language models (LLMs) . Their success is partly due to their versatility in processing diverse data types, including text, images, and video, represented as sequences of tokens (Liu et al., 2021 ; Girdhar et al.,  2019 ; Brown et al., 2020) . While scale has  \n1 Statistical Physics of Computation Laboratory, `Ecole polytechnique fdrale de Lausanne (EPFL), Lausanne, Switzerland 2Department of Computing Sciences, Universit Bocconi, Milan, Italy. Correspondence to: Freya Behrens \u003C[freya.behrens@epfl.ch](freya.behrens@epfl.ch) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nbeen a key factor in unleashing the potential of these models, it is remarkable that their architecture still largely follows the same simple template of the original transformer model proposed by Vaswani et al. (2017) . At its core, a single transformer block primarily alternates two basic components: the token-mixing attention mechanism and a standard fully connected multi-layer perceptron. At a high level, the attention mechanism mixes the tokens, while the multi-layer perceptron applies a nonlinear feature transformation identically to each token. Despite the widespread use of transformers, there is no clear consensus on the distinct roles of their components, how they interact, or if they can be substituted with alternative modules (Tolstikhin et al., 2021 ; Dordevicet al., 2024 ; Gu & Dao, 2024) . In particular, the specific contribution of each architectural element to the model’s hypothesis space –the range of algorithms it can learn and implement in practice– remains opaque (Weiss et al., 2021 ; Deltang et al., 2023 ; Abbe et al., 2023 ; Ouellette et al., 2023) .  \nIn this work, we investigate this question for algorithms that compare and aggregate information from a mechanistic perspective (Cammarata et al., 2020 ; Olah et al., 2020 ; Elhageet al., 2021 ; Michaud et al., 2023 ; Ouellette et al., 2023) and focus on the histogram task (Weiss et al., 2021) as a prototypical problem. It consists of predicting the number of appearances of each token in the input sequence processed by the model – counting-and solving it requires both comparison and aggregation. Even modern language models with up to 8B parameters currently fail to solve this task robustly and efficiently in-weights, rather than via chain-of-thought,(Appendix A) . This failure, despite the task’s apparent simplicity, motivates a study of the relative role of different architectural component","cbCaifRr838X9ZGQ","https://ap.wps.com/l/cbCaifRr838X9ZGQ","pdf",9960793,1,33,"English","en",105,"# Abstract\n# Introduction\n## Transformer architecture components and roles\n## Histogram task and motivation\n## Method overview and contributions","[{\"question\":\"Why can adding softmax and beginning-of-sequence tokens help when embeddings are small?\",\"answer\":\"The study shows these additions improve robustness for comparatively small embedding dimensions, supporting more reliable formation of the learned counting strategies during training.\"}]","Counting in Small Transformers - The Delicate Interplay between Attention and Feed-Forward Layers - Abstract | PDF",1785817111,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"counting-in-small-transformers-the-delicate-interplay-between-attention-and-feed-forward-layers-abstract","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/counting-in-small-transformers-the-delicate-interplay-between-attention-and-feed-forward-layers-abstract/123525/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Why can adding softmax and beginning-of-sequence tokens help when embeddings are small?","Question",{"text":75,"@type":76},"The study shows these additions improve robustness for comparatively small embedding dimensions, supporting more reliable formation of the learned counting strategies during training.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]