[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-150473-en":3,"doc-seo-150473-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},150473,687207017582,"Ethan Miller","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Simulated Attention Score (SAS) - Parameter-Efficient Attention Aggregation (PEAA)","Attention mechanisms are a core component of the Transformer architecture, and performance depends strongly on how attention scores are computed. Building on the observation that increasing the number of attention heads improves results when the per-head hidden size remains sufficiently large, this work introduces Simulated Attention Score (SAS). SAS simulates more heads and higher per-head feature dimensions without increasing parameter count by projecting compact head and embedding representations into larger simulated spaces. To further bound parameter and computation costs, Parameter-Efficient Attention Aggregation (PEAA) is proposed. Experiments across multiple datasets and tasks show consistent gains over attention variants.","arXiv :2507 .07694v2 [ cs .CL] 25 Nov 2025  \nSAS: Simulated Attention Score  \nChuanyang Zheng1 ∗ Jiankai Sun2,3 Yihang Gao4 Yuehao Wang5 Peihao Wang5 Jing Xiong6 Liliang Ren3 Hao Cheng3 Janardhan Kulkarni3 Yelong Shen3 Zhangyang Wang5 Mac Schwager2 Anderson Schneider1 Xiaodong Liu3 Jianfeng Gao3  \n1Morgan Stanley 2 Stanford 3Microsoft Research 4NUS 5UT Austin 6HKU  \nAbstract  \nThe attention mechanism is a core component of the Transformer architecture.  \nVarious methods have been developed to compute attention scores, including multihead attention (MHA), multi-query attention, group-query attention and so on. We further analyze the MHA and observe that its performance improves as the number of attention heads increases, provided the hidden size per head remains sufficiently large. Therefore, increasing both the head count and hidden size per head with minimal parameter overhead can lead to significant performance gains at a low cost.  \nMotivated by this insight, we introduce Simulated Attention Score (SAS), which  \nmaintains a compact model size while simulating a larger number of attention heads and hidden feature dimension per head. This is achieved by projecting a low-dimensional head representation into a higher-dimensional space, effectively increasing attention capacity without increasing parameter count. Beyond the head representations, we further extend the simulation approach to feature dimension of the key and query embeddings, enhancing expressiveness by mimicking the behavior of a larger model while preserving the original model size. To control the parameter cost, we also propose Parameter-Efficient Attention Aggregation (PEAA). Comprehensive experiments on a variety of datasets and tasks demonstrate the effectiveness of the proposed SAS method, achieving significant improvements over different attention variants.  \n1 Introduction  \nThe Transformer architecture [77] has emerged as the predominant framework in modern machine learning, demonstrating remarkable success across diverse domains. Transformer-based models consistently achieve state-of-the-art performance in numerous natural language processing tasks, including machine translation [5, 89], question answering [96, 27, 4], and commonsense reasoning [71, 106, 99, 92, 56] . The fundamental components of the Transformer comprise attention mechanisms and feed-forward networks, with attention score computation serving as a critical element. Beyond the original multi-head attention (MHA) formulation [77], researchers have developed various efficient sparse attention variants, such as sliding window approaches (e.g., Streaming LLMs [80]), linear transformers (e.g., Performer [14]), and sparse attention mechanisms (e.g., Reformer and Sparse Sinkhorn transformer [73, 21]) .  \nModern architectures employ diverse strategies for attention score computation. The standard MHA approach [77] maintains identical head counts across query, key, and value embeddings. To optimize memory efficiency, subsequent innovations have introduced parameter sharing schemes: Multi-Query Attention (MQA) [61] and Group-Query Attention (GQA) [3] share key and value projections across attention heads. The Multi-Latent Attention (MLA) approach [40] adopts low-rank compression for key and value projections, caching only latent representations. More recently, Tensor Product  \n∗ Contact Email: [cyzhengme@gmail.com](cyzhengme@gmail.com)  \n39th Conference on Neural Information Processing Systems (NeurIPS 2025) .  \nAttention (TPA) [100] has introduced tensor decomposition techniques for compact representation of queries, keys, and values. We further analyze the MHA in Figure 1, and we find that the MHA performance gradually increases when the attention head number increases, as long as the hidden size per head is not too small. Hence, by increasing the number of attention heads and the hidden size per head at a minimal cost in terms of additional parameters, we can expect a substantial improvement in ","cbCaiao7Xw6iHCtz","https://ap.wps.com/l/cbCaiao7Xw6iHCtz","pdf",618514,2,1,22,"English","en",105,"# Introduction\n## Transformer attention background and efficient variants\n## Motivation for head and feature simulation\n## Proposed method: SAS and PEAA\n## Contributions and experimental validation","[{\"question\":\"What is the main idea behind Simulated Attention Score (SAS)?\",\"answer\":\"SAS simulates a larger number of attention heads and higher per-head feature dimensions by projecting low-dimensional head representations into higher-dimensional spaces, increasing attention capacity without increasing parameter count.\"},{\"question\":\"How does SAS expand head count and feature dimension for attention score computation?\",\"answer\":\"It applies linear projections with nonlinear activations along the head dimension to expand heads from H to a larger simulated count, and similarly simulates more feature dimensions for query/key embeddings to expand from D to ˆD.\"},{\"question\":\"Why introduce Parameter-Efficient Attention Aggregation (PEAA)?\",\"answer\":\"PEAA is used to control parameter cost and computation size while extending the simulation approach, keeping the overall model compact.\"}]","Simulated Attention Score (SAS) - Parameter-Efficient Attention Aggregation (PEAA) | PDF",1787820925,55,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"simulated-attention-score-sas-parameter-efficient-attention-aggregation-peaa","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/simulated-attention-score-sas-parameter-efficient-attention-aggregation-peaa/150473/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main idea behind Simulated Attention Score (SAS)?","Question",{"text":76,"@type":77},"SAS simulates a larger number of attention heads and higher per-head feature dimensions by projecting low-dimensional head representations into higher-dimensional spaces, increasing attention capacity without increasing parameter count.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SAS expand head count and feature dimension for attention score computation?",{"text":81,"@type":77},"It applies linear projections with nonlinear activations along the head dimension to expand heads from H to a larger simulated count, and similarly simulates more feature dimensions for query/key embeddings to expand from D to ˆD.",{"name":83,"@type":74,"acceptedAnswer":84},"Why introduce Parameter-Efficient Attention Aggregation (PEAA)?",{"text":85,"@type":77},"PEAA is used to control parameter cost and computation size while extending the simulation approach, keeping the overall model compact.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]