[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-150471-en":3,"doc-seo-150471-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},150471,8796095027276,"wps_ap_test_251126_0180","https://avatar.qwps.com/avatar/d3BzX2FwX3Rlc3RfMjUxMTI2XzAxODA=",8,"Research & Report","SAS: Simulated Attention Score - Paper Conference","SAS (Simulated Attention Score) targets Transformer attention-score computation by improving performance without expanding overall model size. The work analyzes multi-head attention and shows gains from increasing head count when per-head hidden size remains sufficiently large. It then projects compact head representations and expands simulated key/query feature dimensions so the attention module behaves like a larger model. Parameter-Efficient Attention Aggregation (PEAA) further controls parameter and computation cost, and experiments across datasets and tasks validate consistent improvements over attention variants.","SAS: Simulated Attention Score  \nChuanyang Zheng1 ∗ Jiankai Sun2,3 Yihang Gao4 Yuehao Wang5 Peihao Wang5 Jing Xiong6 Liliang Ren3 Hao Cheng3 Janardhan Kulkarni3 Yelong Shen3 Zhangyang Wang5 Mac Schwager2 Anderson Schneider1 Xiaodong Liu3 Jianfeng Gao3  \n1Morgan Stanley 2 Stanford 3Microsoft Research 4NUS 5UT Austin 6HKU  \nAbstract  \nThe attention mechanism is a core component of the Transformer architecture.  \nVarious methods have been developed to compute attention scores, including multihead attention (MHA), multi-query attention, group-query attention and so on. We further analyze the MHA and observe that its performance improves as the number of attention heads increases, provided the hidden size per head remains sufficiently large. Therefore, increasing both the head count and hidden size per head with minimal parameter overhead can lead to significant performance gains at a low cost.  \nMotivated by this insight, we introduce Simulated Attention Score (SAS), which  \nmaintains a compact model size while simulating a larger number of attention heads and hidden feature dimension per head. This is achieved by projecting a low-dimensional head representation into a higher-dimensional space, effectively increasing attention capacity without increasing parameter count. Beyond the head representations, we further extend the simulation approach to feature dimension of the key and query embeddings, enhancing expressiveness by mimicking the behavior of a larger model while preserving the original model size. To control the parameter cost, we also propose Parameter-Efficient Attention Aggregation (PEAA). Comprehensive experiments on a variety of datasets and tasks demonstrate the effectiveness of the proposed SAS method, achieving significant improvements over different attention variants.  \n1 Introduction  \nThe Transformer architecture [77] has emerged as the predominant framework in modern machine learning, demonstrating remarkable success across diverse domains. Transformer-based models consistently achieve state-of-the-art performance in numerous natural language processing tasks, including machine translation [5, 89], question answering [96, 27, 4], and commonsense reasoning [71, 106, 99, 92, 56] . The fundamental components of the Transformer comprise attention mechanisms and feed-forward networks, with attention score computation serving as a critical element. Beyond the original multi-head attention (MHA) formulation [77], researchers have developed various efficient sparse attention variants, such as sliding window approaches (e.g., Streaming LLMs [80]), linear transformers (e.g., Performer [14]), and sparse attention mechanisms (e.g., Reformer and Sparse Sinkhorn transformer [73, 21]) .  \nModern architectures employ diverse strategies for attention score computation. The standard MHA approach [77] maintains identical head counts across query, key, and value embeddings. To optimize memory efficiency, subsequent innovations have introduced parameter sharing schemes: Multi-Query Attention (MQA) [61] and Group-Query Attention (GQA) [3] share key and value projections across attention heads. The Multi-Latent Attention (MLA) approach [40] adopts low-rank compression  \n∗ Contact Email: [cyzhengme@gmail.com](cyzhengme@gmail.com)  \n39th Conference on Neural Information Processing Systems (NeurIPS 2025) .  \nfor key and value projections, caching only latent representations. More recently, Tensor Product Attention (TPA) [100] has introduced tensor decomposition techniques for compact representation of queries, keys, and values. We further analyze the MHA in Figure 1, and we find that the MHA performance gradually increases when the attention head number increases, as long as the hidden size per head is not too small. Hence, by increasing the number of attention heads and the hidden size per head at a minimal cost in terms of additional parameters, we can expect a substantial improvement in performance.  \nBuilding on these observations","cbCais9S0GkrgqOW","https://ap.wps.com/l/cbCais9S0GkrgqOW","pdf",539783,1,29,"English","en",105,"# Abstract\n# 1 Introduction\n## Transformer and attention score computation\n## Efficient attention variants\n## Simulated Attention Score (SAS) approach\n## Main contributions","[{\"question\":\"What insight motivates SAS?\",\"answer\":\"The paper analyzes multi-head attention and finds performance improves as the number of attention heads increases, provided the hidden size per head is not too small.\"},{\"question\":\"How does SAS simulate a larger attention mechanism?\",\"answer\":\"SAS projects low-dimensional head representations into higher-dimensional space to simulate more heads and also simulates larger feature dimensions for key/query embeddings, increasing attention capacity without increasing parameter count.\"},{\"question\":\"What role does PEAA play in SAS?\",\"answer\":\"Parameter-Efficient Attention Aggregation (PEAA) is proposed to control parameter cost and computation while using the simulated head and feature expansion strategy.\"}]","SAS: Simulated Attention Score - Paper Conference | PDF",1787820912,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sas-simulated-attention-score-paper-conference","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sas-simulated-attention-score-paper-conference/150471/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What insight motivates SAS?","Question",{"text":75,"@type":76},"The paper analyzes multi-head attention and finds performance improves as the number of attention heads increases, provided the hidden size per head is not too small.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SAS simulate a larger attention mechanism?",{"text":80,"@type":76},"SAS projects low-dimensional head representations into higher-dimensional space to simulate more heads and also simulates larger feature dimensions for key/query embeddings, increasing attention capacity without increasing parameter count.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does PEAA play in SAS?",{"text":84,"@type":76},"Parameter-Efficient Attention Aggregation (PEAA) is proposed to control parameter cost and computation while using the simulated head and feature expansion strategy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]