[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85362-en":3,"doc-seo-85362-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85362,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","An Exact Instrument for State Usage in Selective State-Space Models and the Input-Driven Migration It Reveals","Selective state-space models route information through a bank of first-order modes, with which modes are used determined by a learned selection mechanism tied to the current input. An exact instrument is provided to measure how a trained model uses these modes, exploiting the diagonal state matrix to decompose channel outputs into per-mode contributions and derive closed-form pruning output errors offline for any subset and budget. Validation matches reference and predicts deployed pruning error across thousands of configurations, and shows input-driven migration driven mainly by an input-dependent write map, improving input-scheduled pruning over static baselines.","arXiv :2607 . 1 1796v 1 [ cs .LG] 13 Jul 2026  \nAn Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals  \nRaktim Bhattacharya  \nAerospace Engineering, Texas A&M University  \nAbstract  \nSelective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism. We give an exact instrument for measuring how a trained model uses these modes. Because the state matrix is diagonal, each channel’s output decomposes exactly into per-mode contributions, and a per-(layer, channel, window) Gram tensor yields the exact output error of dropping any subset of modes, offline, at any budget. Validated against the reference implementation to a relative error of 2.3 × 10 −7 on the Mamba-1 family where it is exact, the instrument predicts a layer’s deployed pruning error to a median relative deviation of 5 × 10 −7 over 4 ,464 configurations, its floor set by the reconstruction. Applying the instrument across the Mamba-1 family (130M–2.8B), the deployed 7B Falcon-Mamba, and Mamba-2, we find that trained models re-allocate their state space with the input: which modes carry the signal migrates across contexts, and at the most affected layers a per-input oracle roughly halves the output error of a fixed mode set. Frozen-signal counterfactualsattribute the migration primarily to the input-dependent write map Bt ; the timestep usually identified with selectivity carries almost none of it. Input-scheduled mode pruning on this measurement outperforms static, Hankel-based, and layer-adaptive rankings at every scale from 130M to the deployed 7B Falcon-Mamba, and at half the state budget it matches the unpruned model. Because the scheduler reads each window’s mode usage from a first pass, this demonstrates realizable headroom; we claim no deployed compute or memory saving.  \n1 Introduction  \nSelective state-space models (SSMs) have made linear recurrences competitive with attention for long-sequence modeling (7, 3) . A Mamba layer maintains, in each channel, a small linear state driven by a first-order recurrence with a diagonal state matrix; the layer is selective in that the input, output, and timestep matrices are functions of the current token. The state matrix itself is fixed after training. A layer is therefore a bank of fixed first-order modes whose coupling to the signal is scheduled, token by token, by the selection mechanism.  \nHow a trained model uses that bank is not known. The question matters in two ways. For efficiency, the state is the memory and compute bottleneck of an SSM, and a line of work prunes or reduces it (10, 1); every such method rests on an assumption about which modes matter. For interpretability, the modes are the layer’s internal degrees of freedom, and their usage is a direct measure of what the layer stores. Both uses require an exact measurement of state usage.  \nExisting mode-importance criteria are static and approximate. Output-aware pruning of Mamba-2 scores modes once, from calibration statistics, and prunes a fixed set (10); activity-based criteria rank modes by the timestep (1); the classical reduction they approximate, balanced truncation, is defined for time-invariant systems and bounds a balanced surrogate; the bound does not cover the deployed pruned layer (11, 6) . None measures, per input, which modes a trained layer is using, and none is exact.  \nAn exact measurement is possible because the state matrix is diagonal: the modes are decoupled, and a channel’s scalar output is an exact sum of per-mode contributions. Accumulating the outer products of these contributions over a window yields a small Gram tensor, per (layer, channel, window), from which the exact output error of pruning any subset of modes follows in closed form, offline, at any budget. The decomposition is exact, and we validate it in two ways: against the reference implementation to a relative error of 2 × 10","cbCaiuHpXRumUkwH","https://ap.wps.com/l/cbCaiuHpXRumUkwH","pdf",1440505,1,11,"English","en",105,"# Introduction\n## Selective SSMs and the problem of unknown state usage\n## Limitations of existing mode-importance criteria\n## Exact measurement via diagonal state decomposition\n## Migration across inputs and its quantified impact\n## Causal analysis of migration components\n## Input-scheduled mode pruning results","[{\"question\":\"What does the proposed instrument measure in a selective state-space model?\",\"answer\":\"It measures, exactly, how a trained model uses each first-order mode by decomposing channel outputs into per-mode contributions and deriving the exact pruning output error for any subset of modes offline.\"},{\"question\":\"Why is the measurement exact for these models?\",\"answer\":\"The state matrix is diagonal, so modes are decoupled and a channel output is an exact sum of per-mode contributions; a per-(layer, channel, window) Gram tensor enables closed-form pruning error computation.\"},{\"question\":\"What causes the input-driven migration of mode usage?\",\"answer\":\"Freezing selective signals shows that the input-dependent write map carries most of the migration, the readout carries part, and the timestep (typically associated with selectivity) carries almost none.\"}]",1784202793,28,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"an-exact-instrument-for-state-usage-in-selective-state-space-models-and-the-input-driven-migration-it-reveals","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/an-exact-instrument-for-state-usage-in-selective-state-space-models-and-the-input-driven-migration-it-reveals/85362/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the proposed instrument measure in a selective state-space model?","Question",{"text":75,"@type":76},"It measures, exactly, how a trained model uses each first-order mode by decomposing channel outputs into per-mode contributions and deriving the exact pruning output error for any subset of modes offline.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is the measurement exact for these models?",{"text":80,"@type":76},"The state matrix is diagonal, so modes are decoupled and a channel output is an exact sum of per-mode contributions; a per-(layer, channel, window) Gram tensor enables closed-form pruning error computation.",{"name":82,"@type":73,"acceptedAnswer":83},"What causes the input-driven migration of mode usage?",{"text":84,"@type":76},"Freezing selective signals shows that the input-dependent write map carries most of the migration, the readout carries part, and the timestep (typically associated with selectivity) carries almost none.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]