[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123521-en":3,"doc-seo-123521-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123521,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Fundamental limits of learning in sequence multi-index models and deep attention networks - high-dimensional asymptotics and sharp thresholds","We study learning in deep attention neural networks formed by composing multiple self-attention layers with tied and low-rank weights, and connect them to sequence multi-index models that generalize classical multi-index settings to sequential covariates. Under Bayes-optimal learning with large dimension D and proportionally large sample size N, the work derives a sharp asymptotic characterization of optimal performance, including that of approximate message passing, and identifies sharp thresholds for sample complexity. It further explains how sequential layer learning emerges in realistic setups.","Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds  \nEmanuele Troiani 1 Hugo Cui 2 Yatin Dandi 1 3 Florent Krzakala 3 Lenka Zdeborov 1  \nAbstract  \nIn this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayes-optimal learning, in the limit of large dimension D and proportionally large number of samples N , we derive a sharp asymptotic characterization of the optimal performance as well asthe performance of the best-known polynomialtime algorithm for this setting –namely approximate message-passing–, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup.  \n1. Introduction  \nRecent years have witnessed a shift in paradigm in the automated learning from sequential data, such as language (Brown et al., 2020 ; Kenton & Toutanova, 2019) . The backbone of many of these technological advances arguably lies in the use of transformer architectures (Vaswaniet al., 2017)– a parametrization allowing the model to dynamically focus on relevant portions of the input, and extricate increasingly complex correlations between tokens through the successive application of attention layers. They  \n1 Statistical Physics Of Computation Laboratory, ´Ecole Polytechnique Fdrale de Lausanne (EPFL) 2 Center of Mathematical Sciences and Applications, Harvard University 3Information Learning and Physics Laboratory, ´Ecole Polytechnique Fdrale de Lausanne (EPFL) . Correspondence to: Emanuele Troiani \u003Cemanuele.troiani@epfl.ch, lenka.zdeborova@epfl.ch> .  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nthus constitute a function class able to represent intricate dependencies between the tokens of sequential covariates. In spite of their ubiquity, the study of such models from a theoretical viewpoint is still in its infancy.  \nThis situation stands in contrast with the wealth of resultson multi-index models, central to many theoretical works e.g. (Arous et al., 2021 ; Abbe et al., 2022 ; Veiga et al., 2022 ; Ba et al., 2022 ; Arnaboldi et al., 2023 ; Collins-Woodfinet al., 2024 ; Damian et al., 2024a ; Bietti et al., 2023 ; Moniri et al., 2024 ; Berthier et al., 2024) . Multi-index models define functions based on low-dimensional subspaces ofcovariates and are popular exemplars among theoreticians. However, shallow architectures like multi-index models are structurally simpler than attention models. Specifically, they (a) lack the hierarchical structure from multiple attention layers and (b) act on less structured, non-sequential covariates. These distinctions make it unclear whether results from multi-index models apply to multilayer attention architectures. Here, we show these models share a deep formal connection, enabling a transfer of insights and analytical approaches.  \nNamely, our first motivation is to extend the theoretical framework of multi-index models to sequence models. We consider the class of sequence multi-index (SMI) functions introduced by (Cui et al., 2024a ; Cui, 2025), and defined over length M sequences of D-dimensional tokens x ∈ RD×M of the form  \nySMIW(x) = g 􀀒 xD􀀓 . (1)  \nwhere W ∈ RP×D is a learnable projection matrix, and g : RP×M → RK is a (possibly multi-dimensional) link function. The K−dimensional output could represent –to give an example– the predicted class pro","cbCaifCQ8c7GGqYA","https://ap.wps.com/l/cbCaifCQ8c7GGqYA","pdf",912219,1,36,"English","en",105,"# Abstract\n# Introduction\n## From deep attention to sequence multi-index models\n## Mapping and modeling of architectures\n# Main contributions (Bayes-optimal setting)","[{\"question\":\"What relationship does the paper establish between deep attention networks and sequence multi-index models?\",\"answer\":\"It maps deep attention architectures to sequence multi-index functions, showing a formal connection that transfers theoretical tools and insights between the two model classes.\"},{\"question\":\"In what regime does the paper analyze learning performance and sample complexity?\",\"answer\":\"The analysis targets Bayes-optimal learning when the dimension D is large and the number of samples N grows proportionally, enabling sharp asymptotic results.\"},{\"question\":\"How does the paper characterize learning using approximate message passing?\",\"answer\":\"It derives the asymptotic optimal performance for the setting and also compares it to the performance of approximate message passing, the best-known polynomial-time algorithm for this regime.\"}]","Fundamental limits of learning in sequence multi-index models and deep attention networks - high-dimensional asymptotics and sharp thresholds | PDF",1785817087,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"fundamental-limits-of-learning-in-sequence-multi-index-models-and-deep-attention-networks-high-dimensional-asymptotics-and-sharp-thresholds","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fundamental-limits-of-learning-in-sequence-multi-index-models-and-deep-attention-networks-high-dimensional-asymptotics-and-sharp-thresholds/123521/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What relationship does the paper establish between deep attention networks and sequence multi-index models?","Question",{"text":75,"@type":76},"It maps deep attention architectures to sequence multi-index functions, showing a formal connection that transfers theoretical tools and insights between the two model classes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"In what regime does the paper analyze learning performance and sample complexity?",{"text":80,"@type":76},"The analysis targets Bayes-optimal learning when the dimension D is large and the number of samples N grows proportionally, enabling sharp asymptotic results.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper characterize learning using approximate message passing?",{"text":84,"@type":76},"It derives the asymptotic optimal performance for the setting and also compares it to the performance of approximate message passing, the best-known polynomial-time algorithm for this regime.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]