[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134456-en":3,"doc-seo-134456-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},134456,1099523885074,"Riley West","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","PolySketchFormer - Fast Transformers via Sketching Polynomial Kernels","Self-attention in large-scale Transformer language models incurs quadratic time and memory costs with sequence length, creating a major bottleneck for training and deployment. Recent results show sub-quadratic softmax approximation can be intractable under reasonable assumptions. This work introduces polynomial attention that replaces softmax without degrading model quality and develops linear-time polynomial sketching with approximation guarantees, enabling causal masking efficiently and providing provable end-to-end speedups.","PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels  \nPraneeth Kacham * 1 Vahab Mirrokni * 2 Peilin Zhong * 2  \narXiv :2310 .01655v3 [ cs .LG] 17 Mar 2024  \nAbstract  \nThe quadratic time and memory complexity inherent to self-attention mechanisms, with respect to sequence length, presents a critical computational bottleneck in the training and deployment of largescale Transformer-based language models. Recent theoretical results indicate the intractability of sub-quadratic softmax attention approximation under reasonable complexity assumptions. This paper addresses this challenge by first demonstrating that polynomial attention with high degree can effectively replace softmax without sacrificing model quality. Next, we develop polynomial sketching techniques from numerical linear algebra to achieve linear-time polynomial attention with approximation guarantees. Crucially, our approach achieves this speedup without requiring the sparsification of attention matrices. We also present a block-based algorithm to apply causal masking efficiently. Combining these techniques, we provide PolySketchFormer, a practical lineartime Transformer architecture for language modeling that offers provable guarantees.  \nWe validate PolySketchFormer empirically by training language models capable of handling long contexts. These experiments utilize both synthetic and real-world datasets (PG19, Wikipedia and C4) on Google Cloud TPUs. For context lengths of 32k and GPT-2 style models, our model achieves 2x speedup in training compared to FlashAttention of the fastest configuration, with no observed degradation in quality across our experiments.1  \n1 Carnegie Mellon University 2 Google Research. Correspondence to: Praneeth Kacham \u003C[pkacham@cs.cmu.edu](pkacham@cs.cmu.edu)>, Vahab Mirrokni \u003C[mirrokni@google.com](mirrokni@google.com)>, Peilin Zhong \u003Cpeil[inz@google.com](inz@google.com)>.  \n1Our implementation is available at [https://github](https://github). com/google-research/google-research/tree/master/ polysketchformer  \nTrain Step latency per token  \n2.5 ~~ ~~  \nµs/token  \n2.0  \n1.5  \n1.0  \n0.5  \nVanilla Softmax  \n Polysketch  \n FlashAttention (Block = 256)  \n FlashAttention (Block = 512)  \n0.0 ~~ ~~  \n5000 10000 15000 20000 25000 30000  \nContext Length  \nFigure 1. Train step latency per token in µs/token of GPT-2 small style models with softmax attention (FlashAttention) v.s. ours. Each model is trained with 1M tokens batches. Vanilla softmax attention goes out-of-memory (OOM) for context lengths > 8k.  \n1. Introduction  \nTransformer-based models (Vaswani et al., 2017) are stateof-the-art for many natural language tasks, leading to breakthroughs in machine translation, language understanding (Devlin et al., 2019), and language modeling (Brown et al., 2020 ; Chowdhery et al., 2022 ; OpenAI, 2023 ; Anil et al., 2023) . However, the quadratic time and space complexity of the attention mechanism limits scalability for long context lengths. Numerous “efficient transformers” have been proposed to address this issue (Wang et al., 2020 ; Katharopoulos et al., 2020 ; Choromanski et al., 2020 ; Han et al., 2023) . These variants approximate2 the standard attention mechanism. A survey by Tay et al. (2022) provides a broad overview of these techniques. While many efficient transformer constructions achieve linear theoretical training complexity, the survey observes that practical training speedups are often less significant, with potential losses in model quality. This explains the continued dominance of vanilla transformers.  \nIn this work, we focus on improving training latency for transformer models in decoding-only tasks, specifically language modeling trained via next-word prediction. Our techniques are generalizable to encoding-only and encoderdecoder transformers, with potential applications beyond language modeling. We will first briefly discuss existing approaches to make training of transformer models faster and then place our contri","cbCaif5nIz0hrKWJ","https://ap.wps.com/l/cbCaif5nIz0hrKWJ","pdf",835983,1,21,"English","en",105,"# Introduction\n## Memory efficient and I/O aware approach\n## Approximate softmax attention via sparsification\n## Efficient n×n attention by kernel-based methods","[{\"question\":\"What problem does PolySketchFormer address in Transformer training?\",\"answer\":\"It targets the quadratic time and memory complexity of self-attention with respect to sequence length, which limits scaling to long contexts.\"},{\"question\":\"How does PolySketchFormer avoid the cost of softmax attention?\",\"answer\":\"It replaces softmax with polynomial attention using high-degree polynomials, then uses polynomial sketching techniques to obtain linear-time attention with approximation guarantees.\"},{\"question\":\"What are the reported empirical results for training speed and quality?\",\"answer\":\"On long-context language modeling experiments using PG19, Wikipedia, and C4 on Google Cloud TPUs, PolySketchFormer achieves a 2x training speedup for 32k context length setups compared with the fastest FlashAttention configuration, with no observed quality degradation.\"}]","PolySketchFormer - Fast Transformers via Sketching Polynomial Kernels | PDF",1787262769,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"polysketchformer-fast-transformers-via-sketching-polynomial-kernels","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/polysketchformer-fast-transformers-via-sketching-polynomial-kernels/134456/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does PolySketchFormer address in Transformer training?","Question",{"text":76,"@type":77},"It targets the quadratic time and memory complexity of self-attention with respect to sequence length, which limits scaling to long contexts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does PolySketchFormer avoid the cost of softmax attention?",{"text":81,"@type":77},"It replaces softmax with polynomial attention using high-degree polynomials, then uses polynomial sketching techniques to obtain linear-time attention with approximation guarantees.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the reported empirical results for training speed and quality?",{"text":85,"@type":77},"On long-context language modeling experiments using PG19, Wikipedia, and C4 on Google Cloud TPUs, PolySketchFormer achieves a 2x training speedup for 32k context length setups compared with the fastest FlashAttention configuration, with no observed quality degradation.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]