[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84261-en":3,"doc-seo-84261-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84261,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization","Rotary Position Embeddings (RoPE) give transformers a fixed grid of positional frequencies, but trained models concentrate usage unevenly. The work studies what drives this non-uniform frequency selection and offers a data-centered explanation: frequencies are chosen to match the relative-distance structure of training data. Each frequency is treated as a positional lens, yielding a field-resolution tradeoff where the optimal frequency scales as 1/W for a dependency width W. The principle links RoPE frequency scaling to position-interpolation length generalization. It shows long-context generalization depends on scale matching between learned frequencies and training-time dependencies, and on how those dependencies extend to longer contexts, supported by self-similarity in natural language.","arXiv :2607 .07678v 1 [ cs .LG] 8 Jul 2026  \nHow Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization  \nXinyi Wu∗1 Siyuan Liu∗†2 Ali Jadbabaie 1  \n1MIT IDSS 2IIIS, Tsinghua University  \n{xinyiwu,[jadbabai}@mit.edu](jadbabai}@mit.edu) [liusiyua23@mails.tsinghua.edu.cn](liusiyua23@mails.tsinghua.edu.cn)  \nAbstract  \nRotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly.  \nWe study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution tradeoff and show that, for a data-induced dependency profile of width W , the optimal frequency scales as 1/W. This frequency-matching principle explains controlled observations on synthetic and text-based data, and suggests that the mid-low frequency bands observed in language models arise from the multi-scale dependency structure of natural language. We further connect frequency selection to position-interpolation-based length generalization: scaling frequencies down expands the effective field while reducing resolution. This helps when longer-context dependencies are approximate dilations of those seen during training, but can fail when relevant dependencies do not scale with context length.  \nEmpirically, we show that natural language exhibits approximate self-similarity across positional scales, explaining why test-time frequency scaling can support long-context generalization. Overall, our results identify a data-driven mechanism behind emergent RoPE frequency usage and show that long-context generalization depends on two forms of scale matching: between learned frequencies and trainingtime dependencies, and between frequency scaling and how those dependencies extend to longer contexts.  \n1 Introduction  \nA longstanding view in machine learning is that generalization depends on inductive biases, or preferences that prioritize certain solutions over others when data is finite [4, 28] . One of the central questions in modern machine learning is how such biases arise in models designed to be broadly flexible [34, 43] . Attention is a striking example [2, 20, 39]; although every token can in principle attend to every other token, trained attention models are far from structureless. They exhibit strong positional preferences [14, 22, 43], attention sinks [13, 44], head specialization [30], and highly non-uniform use of embedding dimensions [3, 17, 29] . These phenomena suggest that training does more than fit a predictor: it induces systematic preferences in how a flexible architecture uses its internal degrees of freedom.  \nRotary Position Embeddings (RoPE) provide a particularly clean instance of this phenomenon [33] . RoPE gives a model access to a fixed grid of positional frequencies, yet trained models do not use these frequencies uniformly. Instead, query and key representations often concentrate their norm in restricted frequency bands [3, 17, 29] . This raises a fundamental question:  \nWhat determines which RoPE frequencies a model learns to use? One possible explanation is that high RoPE frequencies encode position while low frequencies  \nencode semantics [3] . This intuition is partly motivated by the fact that very low RoPE frequencies *Equal contribution. † Work done during visit at MIT.  \nPreprint.  \nvary slowly over the training context and are therefore nearly invariant to small changes in position. However, low sensitivity to position should not be mistaken for non-positional information. A low frequency in RoPE still modulates attention scores as a function of relative distance; it simply varies more gradually across the context. Moreover, prior work shows that changing the training sequence length alone can shift RoPE frequency usage [29], suggest","cbCaic7RYdU1UI4a","https://ap.wps.com/l/cbCaic7RYdU1UI4a","pdf",2005747,4,1,21,"English","en",105,"# Abstract\n# Introduction\n## Inductive biases and positional preferences in attention\n## Motivation from non-uniform RoPE frequency usage\n## Data-centered theory: dependency profiles as positional lenses\n## Frequency scaling and the tradeoff between field and resolution\n## Position interpolation and conditions for length generalization","[{\"question\":\"What determines which RoPE frequencies a model uses after training?\",\"answer\":\"The study argues that RoPE frequencies are selected to match the relative-distance structure of the training data, expressed through a data-induced positional dependency profile.\"},{\"question\":\"How is RoPE frequency usage formalized in the proposed theory?\",\"answer\":\"Each RoPE frequency is treated as a positional lens, leading to a field-resolution tradeoff: higher frequencies give sharper local resolution but wrap sooner, while lower frequencies cover broader ranges more coarsely.\"},{\"question\":\"Why does position interpolation help with long-context length generalization, and when can it fail?\",\"answer\":\"It rescales RoPE frequencies so the effective field grows while local resolution decreases, helping when test-time dependencies are approximately stretched versions of training-time ones; it can fail when relevant dependencies do not scale with context length.\"}]",1784194445,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"how-data-shapes-rope-frequency-usage-from-positional-scale-matching-to-length-generalization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/how-data-shapes-rope-frequency-usage-from-positional-scale-matching-to-length-generalization/84261/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What determines which RoPE frequencies a model uses after training?","Question",{"text":75,"@type":76},"The study argues that RoPE frequencies are selected to match the relative-distance structure of the training data, expressed through a data-induced positional dependency profile.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is RoPE frequency usage formalized in the proposed theory?",{"text":80,"@type":76},"Each RoPE frequency is treated as a positional lens, leading to a field-resolution tradeoff: higher frequencies give sharper local resolution but wrap sooner, while lower frequencies cover broader ranges more coarsely.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does position interpolation help with long-context length generalization, and when can it fail?",{"text":84,"@type":76},"It rescales RoPE frequencies so the effective field grows while local resolution decreases, helping when test-time dependencies are approximately stretched versions of training-time ones; it can fail when relevant dependencies do not scale with context length.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]