[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83287-en":3,"doc-seo-83287-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83287,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","The Key to Going Linear Analysis-Driven Transformer Linearization","The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While many post hoc linearization pipelines exist, it remains unclear which components preserve model quality. This work isolates the impact of state update design under a strict frozen-backbone regime, deriving an approximation linking softmax attention to linear state updates. Softmax is shown to apply key-dependent rank-1 orthogonal projections; delta-style updates implement these corrections, while pure gated accumulation misses key geometric structure. Practical interventions including sink tokens, short convolutions, and fixed-budget cache routing reduce approximation error. Experiments across LLaMA and Qwen scale up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching long-context retrieval quality of complex adaptive-caching frameworks.","The Key to Going Linear: Analysis-Driven Transformer Linearization  \nAnna Kuzina  \nQualcomm AI Research∗  \nakuzina@qti.qualcomm.com  \nPaul N. Whatmough  \nQualcomm AI Research  \npwhatmou@qti.qualcomm.com  \nBabak Ehteshami Bejnordi  \nQualcomm AI Research  \nbehtesha@qti.qualcomm.com  \narXiv :2607 .07706v 1 [ cs .LG] 8 Jul 2026  \nAbstract  \nThe quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, short convolutions, and fixed-budget cache routing, which reduces the remaining gap. We scale this linearization approach across LLaMA and Qwen models up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching the long-context retrieval of complex adaptive-caching frameworks.  \n1 Introduction  \nCausal self-attention is the main obstacle to scaling pretrained transformers to longer contexts. Its quadratic compute and growing KV cache make inference increasingly expensive as sequence length grows. A particularly appealing solution is post hoc linearization: converting an existing full-attention model into a linear-time architecture without pretraining from scratch [26, 8, 11, 12, 20, 13] . While recent approaches achieve this, they typically combine multiple interventions simultaneously, such as low-rank adaptation (LoRA), sliding-window attention (SWA), hybrid routing, and distillation. These confounding factors obscure a fundamental design question: which linear state update best approximates pretrained softmax attention?  \nTo answer this, we take an analysis-driven approach to transformer linearization. We isolate the approximation problem by studying attention replacement in a strict regime: the pretrained backbone remains entirely frozen. Only the newly introduced parameters of the replacement mechanism are trained. This setting isolates the capacity of the linear attention replacements.  \nWithin this controlled setting, we derive a first-order approximation connecting softmax attention to linear state updates. Our analysis proves that softmax naturally applies key-dependent, rank-1 orthogonal projections. Delta-style updates (like Gated Delta Networks) natively implement these corrections through rank-1 operators constructed from keys, suppressing components aligned with previous keys. Conversely, pure gated accumulation (like Gated Linear Attention) relies on keyindependent decay, missing this vital geometric correction.  \n∗ Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc.  \nPreprint.  \nWe empirically validate these insights across LLaMA and Qwen backbones. We first compare normalized kernelized linear attention [27], Gated Linear Attention (GLA) [23], and Gated Delta Networks [19] in the strict frozen-backbone setting. Delta-style updates consistently provide the strongest approximation. Guided by a theoretically identified approximation gap, we introduce practical design choices: a disjoint sliding-window path, sink tokens, and projection adaptation through short convolutions or LoRA. We find these components compensate for the limitations of linear mechanisms, narrowing the performance gap and serving as a robust default for post hoc linearization. Our main contributions are as follows:  \n• We study post hoc linearization of pretrained language models in a strict frozen-backbone setting, and empirically compare kernelized, gated, and delta-based linear attention mechanisms as replacements for causal self-attention.  \n• We provide a first-order approximation proving del","cbCaiciIA9obs84x","https://ap.wps.com/l/cbCaiciIA9obs84x","pdf",711705,1,22,"English","en",105,"# Abstract\n# Introduction\n# Background and Related works","[{\"question\":\"What problem does the paper address in long-context transformer inference?\",\"answer\":\"It addresses the quadratic compute of causal self-attention and the resulting growing KV cache cost as sequence length increases, which severely limits long-context inference efficiency.\"},{\"question\":\"How does the paper evaluate which linear state updates best approximate pretrained attention?\",\"answer\":\"It uses an analysis-driven method with a strict frozen-backbone regime, keeping the pretrained transformer entirely frozen and training only the replacement mechanism parameters to isolate the approximation capacity of linear attention updates.\"},{\"question\":\"Why do delta-style updates outperform purely gated accumulation in their analysis?\",\"answer\":\"Softmax is shown to rely on key-dependent rank-1 orthogonal projections. Delta-style (e.g., gated delta) updates naturally implement these corrections through rank-1 operators built from keys, while pure gated accumulation depends on key-independent decay and misses the geometric correction.\"}]",1784186509,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"the-key-to-going-linear-analysis-driven-transformer-linearization","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-key-to-going-linear-analysis-driven-transformer-linearization/83287/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in long-context transformer inference?","Question",{"text":75,"@type":76},"It addresses the quadratic compute of causal self-attention and the resulting growing KV cache cost as sequence length increases, which severely limits long-context inference efficiency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper evaluate which linear state updates best approximate pretrained attention?",{"text":80,"@type":76},"It uses an analysis-driven method with a strict frozen-backbone regime, keeping the pretrained transformer entirely frozen and training only the replacement mechanism parameters to isolate the approximation capacity of linear attention updates.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do delta-style updates outperform purely gated accumulation in their analysis?",{"text":84,"@type":76},"Softmax is shown to rely on key-dependent rank-1 orthogonal projections. Delta-style (e.g., gated delta) updates naturally implement these corrections through rank-1 operators built from keys, while pure gated accumulation depends on key-independent decay and misses the geometric correction.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]