[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86341-en":3,"doc-seo-86341-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86341,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Requential Coding Pushing the Limits of Model Compression with Self Generated Training Data","Compression is presented as a core driver of generalization: short codes can capture training regularities, yet parameter based compression often scales with model size and trajectory based schemes can grow with data entropy. Requential coding addresses this by using a teacher model to select student training samples from the student’s own distribution and recording only disagreement indices. The resulting code length is independent of parameter count and data entropy, often far shorter than prequential coding, improving with scale and enabling tighter PAC-Bayes generalization guarantees and analysis of overfitting, learnable structure, and entropy effects.","Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data  \nShikai Qiu  \nNew York University  \nMarc Finzi  \nCarnegie Mellon University  \nYujia Zheng  \nCarnegie Mellon University  \narXiv :2607 . 1 1883v 1 [ cs .LG] 13 Jul 2026  \nKun Zhang  \nCarnegie Mellon University  \nAndrew Gordon Wilson  \nNew York University  \nAbstract  \nCompression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functions far simpler than their parameter counts suggest, but it is challenging to construct codes that realize this simplicity. Parameter-based methods such as quantization produce code lengths that scale with model size, insensitive to how much information the parameters store. Prequential coding bypasses this issue by compressing the training trajectory, but codes the exact data sequence regardless of how much the model learns, yielding large codes when the data has high entropy. We introduce requential coding†, where a teacher model selects training samples drawn from the student’s own distribution. The student’s code records only these selections, which cost bits only where teacher and student disagree. The resulting code length is independent of parameter count and data entropy, and often orders of magnitude shorter than the prequential counterpart, with an advantage that grows with scale. This compression sheds light on phenomena inaccessible to prior compressors. Holding loss fixed, larger models and ensembles compress to much smaller sizes despite more parameters. Plugged into a PAC-Bayes bound, the requential code yields state-of-the-art generalization guarantees for billion-parameter LLMs, outperforming bounds built on aggressive post-training quantization even granted zero error. The bound tightens with scale in the compute-optimal regime, as models become increasingly compressible relative to dataset size. The same code predicts that models gradually overfit when trained for multiple epochs. It also isolates the learnable information in a dataset from its unpredictable, random content, revealing that lower-entropy text holds far more learnable structure than higher-entropy image data.  \n1 Introduction  \nMeasuring compression is key to understanding generalization in deep learning. In order to compress data, a model must discover regularities that facilitate generalization. This intuition underlies fundamental principles of induction, such as Occam’s razor: the simplest explanation consistent with observations is most likely to be true. A strong enough compression can guarantee a model’s generalization performance, limit memorization, and even reveal how much learnable information content is in the training data. Indeed, a growing body of evidence suggests that neural networks often learn functions far simpler than their parameters could express [27, 10, 53, 30, 48] .  \nHowever, finding a sufficiently good compression at scale remains a fundamental open question. It could be that larger neural networks find even simpler, more compressible functions, but demonstrating this compressibility becomes increasingly difficult with scale. As we scale model and data size, existing model compression schemes are inflated by quantities unrelated to actual learning:  \n†Code available at [https://github.com/shikaiqiu/requential-coding](https://github.com/shikaiqiu/requential-coding).  \n0 2 4 6 8 10  \nTokens (B)  \n107 108 109  \nParameters  \nFigure 1: Requential coding achieves strong model compression. (Left) The student model Pt being compressed samples candidates Yt(0) , Yt(1) ,   i..d. Pt for its own training data. A teacher model Qt acceptsan index i⋆t chosen so that Xt = Yt(i⋆t) is marginally distributed as Qt , and the student trains on Xt to yield Pt+1 . A message mt of about KL(Qt ∥Pt) bits encodes the accepted index using relative entropy coding (REC) and is appended to the stud","cbCaiqwyD1efr16Q","https://ap.wps.com/l/cbCaiqwyD1efr16Q","pdf",891988,5,1,24,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does requential coding address in existing model compression methods?\",\"answer\":\"It addresses two limitations: parameter based quantization yields code lengths that scale with model size, and prequential coding yields codes that scale with dataset size by encoding the exact training sequence regardless of how much is learned.\"},{\"question\":\"How does requential coding generate the compressed code?\",\"answer\":\"A teacher model selects training samples drawn from the student’s own distribution, and the student’s code records only the accepted indices. Bits are spent mainly where teacher and student disagree, using disagreement costs such as KL-based relative entropy coding.\"},{\"question\":\"What benefits does requential coding provide for generalization and analysis?\",\"answer\":\"With a PAC-Bayes bound, the requential code yields state of the art generalization guarantees for billion parameter LLMs and tightens with scale in the compute optimal regime. It also predicts gradual overfitting and helps separate learnable information from unpredictable content, showing lower entropy text holds more learnable structure than higher entropy image data.\"}]",1784210575,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"requential-coding-pushing-the-limits-of-model-compression-with-self-generated-training-data","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/requential-coding-pushing-the-limits-of-model-compression-with-self-generated-training-data/86341/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does requential coding address in existing model compression methods?","Question",{"text":76,"@type":77},"It addresses two limitations: parameter based quantization yields code lengths that scale with model size, and prequential coding yields codes that scale with dataset size by encoding the exact training sequence regardless of how much is learned.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does requential coding generate the compressed code?",{"text":81,"@type":77},"A teacher model selects training samples drawn from the student’s own distribution, and the student’s code records only the accepted indices. Bits are spent mainly where teacher and student disagree, using disagreement costs such as KL-based relative entropy coding.",{"name":83,"@type":74,"acceptedAnswer":84},"What benefits does requential coding provide for generalization and analysis?",{"text":85,"@type":77},"With a PAC-Bayes bound, the requential code yields state of the art generalization guarantees for billion parameter LLMs and tightens with scale in the compute optimal regime. It also predicts gradual overfitting and helps separate learnable information from unpredictable content, showing lower entropy text holds more learnable structure than higher entropy image data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]