[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84574-en":3,"doc-seo-84574-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84574,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization","Transformer-based large language models leverage context-aware attention to achieve strong in-context learning without parameter updates, yet softmax attention incurs quadratic computational and memory cost as context length grows, limiting long-context capability. Linear transformers reduce complexity to linear dependence on context length, but their feature-mapping theory and generalization behavior remain unclear. This paper studies approximation and generalization of linear transformers using a domain-generalization view and introduces activation and loss design guidance for linearized pretrained softmax models.","arXiv :2607 .00479v 1 [ cs .LG] 1 Jul 2026  \nGhost in the Kernel: In-Context Learning with Eﬀicient Transformers via Domain Generalization  \nPeilin Liu 1 􀀃 Ding-Xuan Zhou 1y  \n1 School of Mathematics and Statistics, University of Sydney, NSW Australia  \nAbstract  \nTransformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning. With richer context, transformers adapt more effectively to the current use case without any parameter updates. However, the quadratic computational and memory complexity with respect to context length significantly slows data processing in softmax transformers. Linear transformers were proposed to address this issue by reducing the complexity to linear dependence on context length, but the design and understanding of the feature mapping in linear attention, from a theoretical viewpoint, remain unclear. In this paper, we investigate the approximation and generalization abilities of linear transformers under a two-staged sampling process from domain generalization.  \nWe show that linear transformers perform in-context learning as learning a mapping from context distributions to response functions. A dimension-independent convergence rate is obtained for our generalization analysis, which also exhibits the tradeoff between the regularities of data distributions and latent features. Guided by our theoretical framework, we propose a new perspective on activation and loss design for linearizing pretrained softmax large language models.  \nKeywords: in-context learning, operator learning, eﬀicient transformer, linear attention, generalization analysis  \n1 Introduction  \nTransformer-based neural networks have become the foundation of modern deep learning frameworks for natural language processing [6, 12] and computer vision [17, 31] . Especially when trained with large and diverse corpora, transformers exhibit remarkable few/zero-shot generalization capabilities across various downstream tasks [6] . An underlying mechanism for this behavior is in-context learning, in which pretrained large language models (LLMs) condition on instructions or a few input-output pairs (both referred to as prompts) and make predictions on test examples without parameter updates. Many theoretical and empirical studies [1, 15 , 48 , 55] have demonstrated that the emergence of in-context learning capability is closely related with the context-aware structure of the attention mechanism [see 46] which enables each token in a sequence to adaptively weight information from all other tokens and to produce a representation conditioned on the sequence context.  \nHowever, the original softmax attention mechanism in Vaswani et al. [46] is a double-edged sword: while it is beneficial for context-based representation learning, it suffers from quadratic computational complexity with respect to the context length [see 52] . As context length grows dramatically, the quadratic computational and memory costs of the standard attention increasingly hinder autoregressive training and inference, undermining LLM performance in long-context modeling scenarios  \n∗. Email: [peilin.6liu@gmail.com](peilin.6liu@gmail.com)  \n†. Email: [dingxuan.zhou@sydney.edu.au](dingxuan.zhou@sydney.edu.au)  \nsuch as processing entire codebases, preserving coherence in long conversations and performing indepth reasoning across several documents. Therefore, it’s crucial for the design of LLMs to alleviate the curse of quadratic complexity and improve long-context processing capabilities. Numerous works have been recently proposed with this motivation and show performance comparable to the standard attention, including RetNet [43], Mamba [16], and Gated Linear Attention [50] . One branch of these works is known as linear attention [see 7 , 20 , 50 , 54], which replaces the exponential similarity function with a dot product of key/query functions and yields a linear ","cbCaibXL4NHATFlW","https://ap.wps.com/l/cbCaibXL4NHATFlW","pdf",967230,1,53,"English","en",105,"# Abstract\n# Introduction\n# Problem Motivation\n# Background and Related Work\n# Method Overview and Contributions","[{\"question\":\"What problem does the paper address in transformer-based in-context learning?\",\"answer\":\"It addresses the quadratic computational and memory complexity of softmax attention with respect to context length, which slows processing and weakens performance in long-context scenarios.\"},{\"question\":\"How does the paper connect linear transformers to domain generalization?\",\"answer\":\"It shows that in-context learning can be interpreted through a domain-generalization framework, where transformer behavior maps context distributions to response functions.\"},{\"question\":\"What is the main contribution of the theoretical framework?\",\"answer\":\"It provides an analysis framework for linear transformers under a two-staged sampling process, yielding a dimension-independent convergence rate and revealing tradeoffs between distribution regularities and latent features.\"}]",1784196888,134,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"ghost-in-the-kernel-in-context-learning-with-efficient-transformers-via-domain-generalization","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/ghost-in-the-kernel-in-context-learning-with-efficient-transformers-via-domain-generalization/84574/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in transformer-based in-context learning?","Question",{"text":75,"@type":76},"It addresses the quadratic computational and memory complexity of softmax attention with respect to context length, which slows processing and weakens performance in long-context scenarios.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper connect linear transformers to domain generalization?",{"text":80,"@type":76},"It shows that in-context learning can be interpreted through a domain-generalization framework, where transformer behavior maps context distributions to response functions.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main contribution of the theoretical framework?",{"text":84,"@type":76},"It provides an analysis framework for linear transformers under a two-staged sampling process, yielding a dimension-independent convergence rate and revealing tradeoffs between distribution regularities and latent features.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]