[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81605-en":3,"doc-seo-81605-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81605,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Lost in Backpropagation: The LM Head is a Gradient Bottleneck","Last-layer neural language models map hidden features of dimension D to vocabulary-size logits in dimension V, with D much smaller than V, creating a “softmax bottleneck.” The work shows this mismatch limits both expressivity and optimization: backpropagating V-dimensional gradients through a rank-D linear LM head induces unavoidable gradient compression. Empirical results suppress 95–99% of gradient norm, yielding suboptimal update directions and altered training dynamics, including inability to learn even trivial patterns.","arXiv :2603 . 10 145v2 [ cs .CL] 10 Jul 2026  \nLost in Backpropagation:  \nThe LM Head is a Gradient Bottleneck  \nNathan Godey, Yoav Artzi  \nCornell University New York, USA [godeynathan@gmail.com](godeynathan@gmail.com)  \nAbstract  \nThe last layer of neural language models (LMs) projects output features of dimension D to logits in dimension V, the size of the vocabulary, where usually D ≪ V. This mismatch is known to raise risks of limited expressivity in neural LMs, creating a so-called softmax bottleneck. We show the softmax bottleneck is not only an expressivity bottleneck but also an optimization bottleneck. Backpropagating V-dimensional gradients through a rank-D linear layer induces unavoidable compression, which alters the training feedback provided to the vast majority of the parameters. We present a theoretical analysis of this phenomenon and measure empirically that 95- 99% of the gradient norm is suppressed by the output layer, resulting in vastly suboptimal update directions. We conduct controlled pretraining experiments showing that the gradient bottleneck makes trivial patterns unlearnable, and drastically affects the training dynamics of LLMs. We argue that this inherent flaw contributes to training inefficiencies at scale independently of the model architecture, and raises the need for new LM head designs.  \n1 Introduction  \nThere is significant research effort focusing on developing new architectures for the hidden layers of language models (LMs), aiming to improve efficiency and performance at training and inference times (e.g. Gu & Dao, 2024; Ye et al., 2025; Yang et al., 2025) . Despite fundamental differences, all considered architectures share a common structure for the output layer: a single linear mapping, followed by a softmax. This component is often called the LM head. In other words, most autoregressive LMs can be seen as (potentially massive) feature extractors for a multi-class classifier, where classes correspond to tokens.  \nThe LM head design is a standard choice for multi-class classification. However, the language modeling classification setup itself is fairly unusual: the number of extracted features, i.e. the hidden dimension D of the model, is commonly orders of magnitude smaller than the number of classes, i.e. the token vocabulary size V. This mismatch has been shown to limit the expressivity of LMs, and may lead to representation degeneration and performance saturation for small models (Yang et al., 2018; Ganea et al., 2019; Godeyet al., 2024) . Yet, an important aspect of this mismatched mechanism has so far been ignored: the loss gradients in the high-dimensional logit space are backpropagated to a lower-dimension space before the backward pass reaches the rest of the layers. In this work, we frame the softmax bottleneck not as an expressivity issue but instead as a significant factor in training dynamics, where stronger bottlenecks hurt the data efficiency of LMs regardless of the underlying architectural choices.  \nWe show both empirically and theoretically that the softmax bottleneck induces lossy compression during backpropagation, destroying >95% of the gradient norm. Low-rank LM heads hamper optimization dynamics, making some trivial patterns unlearnable (Section 3.2), and reducing LLM training efficiency by up to ×16 for the same backbone (Sec-  \n(a) Expressivity view  \n(b) Optimization view  \nFigure 1: Two views of the softmax bottleneck. (a) Expressivity view: The LM head Wθ projects hidden states Hθinto a rank-D subspace of the V-dimensional log-probability space, constraining where model log-probabilities (•) can lie relative to the true log-probabilities (•); (b) Optimization view: During backpropagation, the full logit gradient (→) lives in the V-dimensional space, but only a projected component (→) is visible when passing through Wθ back to the hidden state space.  \ntion 3.1) . Our work sheds light on a so far overlooked weakness in current LLM design,  \noutlining import","cbCaiijhqcgJHwd9","https://ap.wps.com/l/cbCaiijhqcgJHwd9","pdf",1691263,2,1,29,"English","en",105,"# Abstract\n# Introduction\n## Expressivity view\n## Optimization view\n# Theoretical Overview\n## Problem Setup","[{\"question\":\"What is the “softmax bottleneck” in neural language models?\",\"answer\":\"It is the mismatch where a low hidden dimension D is mapped to a much higher vocabulary-sized logit dimension V in the LM head, typically using a linear layer plus softmax.\"},{\"question\":\"How does the bottleneck affect training beyond expressivity?\",\"answer\":\"Backpropagating gradients through the low-rank LM head compresses and suppresses most of the gradient norm, distorting the optimization signal for earlier layers.\"},{\"question\":\"What empirical evidence supports the gradient bottleneck claim?\",\"answer\":\"The document reports that 95–99% of the gradient norm is suppressed by the output layer, leading to vastly suboptimal update directions and reduced learning efficiency.\"}]",1784174719,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lost-in-backpropagation-the-lm-head-is-a-gradient-bottleneck","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/lost-in-backpropagation-the-lm-head-is-a-gradient-bottleneck/81605/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the “softmax bottleneck” in neural language models?","Question",{"text":75,"@type":76},"It is the mismatch where a low hidden dimension D is mapped to a much higher vocabulary-sized logit dimension V in the LM head, typically using a linear layer plus softmax.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the bottleneck affect training beyond expressivity?",{"text":80,"@type":76},"Backpropagating gradients through the low-rank LM head compresses and suppresses most of the gradient norm, distorting the optimization signal for earlier layers.",{"name":82,"@type":73,"acceptedAnswer":83},"What empirical evidence supports the gradient bottleneck claim?",{"text":84,"@type":76},"The document reports that 95–99% of the gradient norm is suppressed by the output layer, leading to vastly suboptimal update directions and reduced learning efficiency.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]