[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82107-en":3,"doc-seo-82107-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82107,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Training, Reading, and Editing Legible Transformers","A transformer can be constructed from operators that are legible by design, using bounded, named units interpreted as fuzzy set operations rather than dense activations, but legibility requires sustained training pressure. Over-sharpening via a crispness penalty can fail by collapsing bounded operators into dead constants. Using an identity relating crispness, mean, and variance, the work introduces a per-channel variance floor loss that distinguishes live contextual detectors from constants, restoring both legibility and output quality. A learned gating strategy routes 87% of load through crisp operators.","arXiv :2607 .08946v 1 [ cs .LG] 9 Jul 2026  \nTraining, Reading, and Editing Legible Transformers  \nMark Oskin  \nProfessor  \nSchool of Computer Science and Engineering  \nUniversity of Washington  \n[mhoskin@uw. edu](mhoskin@uw. edu)  \nJuly 2026  \nAbstract  \nA transformer can be built from operators that are legible by construction—bounded, named units that read as fuzzy set operations rather than dense activations—but legibility must be pressed for during training, and the pressure has a failure mode. A crispness penalty meant to sharpen a bounded operator into a decisive detector instead collapses it into a dead constant. An identity, E[v(1−v)] = µ(1−µ) − var, shows why—the penalty is a variance-minimizer blind to the difference between a live detector and a constant—and names the fix: a per-channel variance floor, the target legibility metric written as a loss, which recovers both legibility and quality. A learned per-unit fraction then retires the hand-set reservedGELU partition of prior work: given the choice the model keeps no unit as pure GELU and routes 87% of its load-bearing computation through crisp operators. The result is the most legible transformer we have built — 78% of its feed-forward operands and 50% of its attention value channels are crisp-andcontextual detectors, and per-head legibility rises from 18% in shallow layers to 78% in deep ones. Read in the correct rotated per-layer frame, these units separate a clean detection (what a unit responds to) from a harder naming (what its output decodes to); and because the objective makes each unit crisp and sparse, edits to them are far more local—50–184× in the deep layers where the edit sites concentrate—and can target explicit conjunctions a single neuron cannot express. Finally, a between-unit decorrelation pressure exposes a legibility dial: it trades a circuit’s reuse for independence at no quality cost, turning concepts into single, surgically editable units and a prediction into a short explanation read off a handful of named operations. Quality holds at parity with a conventional baseline throughout.  \n1 Introduction  \nMost interpretability reconstructs meaning after training: a dense activation is probed, decomposed with a sparse autoencoder, or decoded through the vocabulary, and a human names what was found. An alternative is to build the meaning in—to make a model’s operators legible by construction, so that a coordinate carries a fixed, stated reading rather than one recovered post hoc. Two prior constructions take this route: the feed-forward layer becomes fuzzy set operations on bounded operands (Oskin, 2026b), and the attention value becomes bounded, named detectors (Oskin, 2026a), each at language-model parity (Figure 1) . But a legible-by-construction substrate is not yet a legible model. Its operators read cleanly only if training keeps them so; the two constructions have never been trained together; and reading such a model, and editing it, are problems of their own. This paper is about making that line usable end to end—how to train the model so its units are actually legible, how to read them, how to edit them, and how a single further pressure reshapes the circuits they form.  \nThe obstacle is the very pressure legibility depends on. A bounded operator reads cleanly only when it fires decisively—near 0 or near 1,“absent” or “present”—so both constructions add a crispness penalty that pushes each value to a rail. Pushed too hard, that penalty does the opposite of its purpose: it drives a unit toa fixed value that ignores its input, a dead constant, perfectly crisp and carrying no information, and quality collapses with it. The cause is exact. For a bounded value the crispness penalty equals µ(1 − µ) − var, a variance-minimizer that cannot tell a live contextual detector from a dead constant, so gradient descent takes  \nthe cheaper constant. Naming the cause gives the cure: a per-channel variance floor — the property that separates a live detect","cbCaikDQAl62sczv","https://ap.wps.com/l/cbCaikDQAl62sczv","pdf",2007376,1,33,"English","en",105,"# Introduction\n## Legible-by-construction operators\n## Failure mode of crispness training pressure\n## Variance-floor loss for recoverable legibility\n## Reading legible models with per-layer tuned lenses\n## Editing benefits from crisp sparse units","[{\"question\":\"Why does a crispness penalty sometimes collapse legible operators into dead constants?\",\"answer\":\"For bounded values, the crispness penalty becomes a variance-minimizer that cannot distinguish a live contextual detector from an input-ignoring constant, so optimization selects the cheaper constant and the quality degrades.\"},{\"question\":\"What is the proposed fix to recover both legibility and quality during training?\",\"answer\":\"The method introduces a per-channel variance floor, implemented as an additional loss term, which separates live detectors from dead constants and restores the intended crisp, informative behavior.\"},{\"question\":\"How does the paper improve legibility coverage across transformer sublayers?\",\"answer\":\"A learned per-unit fraction replaces a fixed GELU reservation: the model avoids keeping any unit as pure GELU and routes most load-bearing computation through crisp operators, achieving high proportions of crisp feed-forward operands and attention value channels.\"}]",1784178244,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"training-reading-and-editing-legible-transformers","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/training-reading-and-editing-legible-transformers/82107/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does a crispness penalty sometimes collapse legible operators into dead constants?","Question",{"text":75,"@type":76},"For bounded values, the crispness penalty becomes a variance-minimizer that cannot distinguish a live contextual detector from an input-ignoring constant, so optimization selects the cheaper constant and the quality degrades.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed fix to recover both legibility and quality during training?",{"text":80,"@type":76},"The method introduces a per-channel variance floor, implemented as an additional loss term, which separates live detectors from dead constants and restores the intended crisp, informative behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper improve legibility coverage across transformer sublayers?",{"text":84,"@type":76},"A learned per-unit fraction replaces a fixed GELU reservation: the model avoids keeping any unit as pure GELU and routes most load-bearing computation through crisp operators, achieving high proportions of crisp feed-forward operands and attention value channels.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]