[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82525-en":3,"doc-seo-82525-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82525,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Prototype Language Models","Knowing which training examples drive model outputs is crucial for auditing, correcting, and understanding modern LLM behavior, yet existing language-model explanations remain expensive, approximate, and largely post-hoc. This work proposes PRISM, a prototype language-model architecture that predicts via a sparse, non-negative mixture of learned prototypes. Clustering objectives anchor prototypes to coherent neighborhoods in training data. Across 130M–1.6B parameters and up to 50B tokens, PRISM matches or exceeds dense baselines, localizes curvature for tractable Hessians, and enables fast attribution (~500×).","arXiv :2607 .005 10v 1 [ cs .LG] 1 Jul 2026  \nPrototype Language Models  \nDan Ley1,2* Giang Nguyen2 Himabindu Lakkaraju1 Julius Adebayo2  \n1 Harvard University 2 Guide Labs Inc.  \n*  \nWork initiated during an internship at Guide Labs.  \nAbstract   \nKnowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models generate tokens through a dense network pathway, causing training data's influence to be distributed across parameters rather than organized along explicit, traceable components. We introduce a prototype language model architecture, Prototypes for Interpretable Sequence Modeling (PRISM), that forms each prediction via a sparse, non-negative mixture of learned prototypes, trained with clustering objectives that anchor each prototype to coherent neighborhoods of training examples. Across architectures from 130M to 1.6B parameters trained on up to 50B tokens, prototype language models either surpass or remain within 2.5 percentage points on average downstream accuracy of matched dense baselines. We show that sparse prototype structure localizes curvature in the loss landscape, yielding a more tractable Hessian and enabling training data attribution that is ∼500 × faster than post hoc baselines when consuming equivalent memory. Calibrating linear prototype controllers can improve downstream accuracy by roughly 3 points while tracing those corrections back to training neighborhoods, and targeted prototype suppression can remove model behaviors without finetuning or measurable loss in generation quality.  \nStandard language models entangle training effects, while PRISM makes them traceable through prototypes.  \nContents  \n1 Introduction 1  \n2 PRISM: Prototypes for Interpretable Sequence Modeling 3  \n2.1 Warm-up: reviving ProtoPNet for language modeling ..................... 3  \n2.2 Sparse prototypical reconstruction ................................. 4  \n2.3 Automated interpretability pipeline ................................. 6  \n3 Using TinyStories as a Microscope 6  \n3.1 Anatomy of a prototype prediction ................................. 7  \n3.2 Training and interpretability dynamics ............................... 8  \n3.3 Prototype structure is learned into the backbone ........................ 10  \n4 Training Data Attribution in a Prototype Subspace 11  \n4.1 Clustering localizes curvature in prototype space ........................ 12  \n4.2 Cacheable prototype space influence functions ......................... 14  \n4.3 Cached attribution preserves signal at lower cost ........................ 15  \n5 Scaling Prototype Language Models 17  \n5.1 Scaling setup and downstream performance ........................... 17  \n5.2 Training stability and efficiency .................................... 19  \n6 New Workflows Enabled by Prototypes 19  \n6.1 Understanding model behavior through prototypes ...................... 20  \n6.2 Prototype controllers boost performance and trace corrections to training data .... 21  \n6.3 Preference alignment without finetuning ............................. 23  \n7 Limitations and Future Roadmap 24  \n7.1 Sequence and document attribution ................................ 25  \n7.2 Hessian-aware model design ..................................... 25  \n7.3 Deep prototype language models .................................. 25  \n7.4 Retraining and learned attribution objectives ........................... 26  \n8 Related work 26  \n9 Conclusion 28  \n1 Introduction  \nWhen a language model (LM) produces a harmful response, copyrighted content, or makes an inaccurate factual claim, a natural follow-up question is: which training data made that output likely? Training data attribution (TDA) consists of a family of techniques to answer this question, characterizing how individual examples, groups of examples, or broader training sources influence model predic","cbCaiej6hOhYqE9V","https://ap.wps.com/l/cbCaiej6hOhYqE9V","pdf",10021384,2,1,54,"English","en",105,"# Introduction\n# PRISM: Prototypes for Interpretable Sequence Modeling\n## Warm-up: reviving ProtoPNet for language modeling\n## Sparse prototypical reconstruction\n## Automated interpretability pipeline\n# Using TinyStories as a Microscope\n## Anatomy of a prototype prediction\n## Training and interpretability dynamics\n## Prototype structure is learned into the backbone\n# Training Data Attribution in a Prototype Subspace\n## Clustering localizes curvature in prototype space\n## Cacheable prototype space influence functions\n## Cached attribution preserves signal at lower cost\n# Scaling Prototype Language Models\n## Scaling setup and downstream performance\n## Training stability and efficiency\n# New Workflows Enabled by Prototypes\n## Understanding model behavior through prototypes\n## Prototype controllers boost performance and trace corrections to training data\n## Preference alignment without finetuning\n# Limitations and Future Roadmap\n## Sequence and document attribution\n## Hessian-aware model design\n## Deep prototype language models\n## Retraining and learned attribution objectives\n# Related work\n# Conclusion","[{\"question\":\"What problem does the paper target in language model interpretability?\",\"answer\":\"It targets training data attribution: identifying which training examples most likely caused a specific model output, despite standard LLMs entangling effects across many parameters and making direct tracing difficult.\"},{\"question\":\"How does PRISM make predictions and improve interpretability?\",\"answer\":\"PRISM forms each next-token prediction via a sparse, non-negative mixture of learned prototypes. Prototype clustering objectives anchor each prototype to coherent neighborhoods of training examples, making effects traceable through the prototype structure.\"},{\"question\":\"What evidence is given that prototype models are effective and efficient?\",\"answer\":\"Across model sizes (130M to 1.6B) and up to 50B tokens, prototype language models either match or improve downstream accuracy versus dense baselines by about 2.5 percentage points. The sparse structure localizes curvature, yields a more tractable Hessian, and enables training-data attribution roughly 500× faster than post-hoc baselines at comparable memory use.\"}]",1784181242,136,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prototype-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/prototype-language-models/82525/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper target in language model interpretability?","Question",{"text":75,"@type":76},"It targets training data attribution: identifying which training examples most likely caused a specific model output, despite standard LLMs entangling effects across many parameters and making direct tracing difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does PRISM make predictions and improve interpretability?",{"text":80,"@type":76},"PRISM forms each next-token prediction via a sparse, non-negative mixture of learned prototypes. Prototype clustering objectives anchor each prototype to coherent neighborhoods of training examples, making effects traceable through the prototype structure.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is given that prototype models are effective and efficient?",{"text":84,"@type":76},"Across model sizes (130M to 1.6B) and up to 50B tokens, prototype language models either match or improve downstream accuracy versus dense baselines by about 2.5 percentage points. The sparse structure localizes curvature, yields a more tractable Hessian, and enables training-data attribution roughly 500× faster than post-hoc baselines at comparable memory use.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]