[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86257-en":3,"doc-seo-86257-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86257,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Random Label Prediction Heads for Studying Memorization in Deep Neural Networks","A straightforward method is introduced for empirically studying memorization in deep neural networks on classification tasks. Each training sample is augmented with auxiliary random labels, then predicted by a Random Label Prediction (RLP) head attached at arbitrary network depths using intermediate representations. Treating RLP-head accuracy as an empirical estimate of Rademacher complexity yields direct measures of sample-level memorization and model capacity. The approach analyzes generalization and overfitting across models and datasets and proposes an RLP-based regularizer that reduces memorization, with effects on generalization that vary by setup.","arXiv :2607 . 1 154 1v 1 [ cs .LG] 13 Jul 2026  \nRANDOM LABEL PREDICTION HEADS FOR STUDYING MEMORIZATION IN DEEP NEURAL NETWORKS  \nMarlon Becker Jonas Konrad Luis Garcia Rodriguez Benjamin Risse  \nUniversity of M¨unster, Germany  \n{marlonbecker,jonas.konrad,luis.garcia,[b.risse](b.risse}@uni-muenster.de)[}](b.risse}@uni-muenster.de)[@uni-muenster.de](b.risse}@uni-muenster.de)  \nABSTRACT  \nWe introduce a straightforward yet effective method to empirically study memorization in deep neural networks for classification tasks. Our approach augments each training sample with auxiliary random labels, which are then predicted by a random label prediction head (RLP-head) . RLP-heads can be attached at arbitrary depths of a network, predicting random labels from the corresponding intermediate representation and thereby enabling analysis of how memorization capacity evolves across layers. By interpreting the RLP-head performance asan empirical estimate of Rademacher complexity, we obtain a direct measure of both sample-level memorization and model capacity. We leverage this random label accuracy metric to analyze generalization and overfitting in different models and datasets. Building on this approach, we further propose a novel regularization technique based on the output of the RLP-head, which demonstrably reduces memorization. Interestingly, our experiments reveal that reducing memorization can either improve or impair generalization, depending on the dataset and training setup. These findings challenge the traditional assumption that overfitting is equivalent to memorization and suggest new hypotheses to reconcile these seemingly contradictory results. The source code is available at [https://github.com/MarlonBecker/RandomLabelHeads](https://github.com/MarlonBecker/RandomLabelHeads).  \n1 INTRODUCTION  \nModern deep learning models are prone to overfitting due to their extreme over-parameterization (Nakkiran et al., 2021) . A wide range of strategies have been proposed to mitigate this issue, including data augmentation, explicit regularization, and dataset scaling. Although enlarging training datasets has proven particularly effective, this approach is often infeasible in domains where data acquisition or annotation is expensive or requires significant human expertise. Moreover, existing strategies primarily address practical concerns of generalization but provide limited insight into the mechanisms by which overfitting arises.  \nRecent work highlights the striking memorization capacity of state-of-the-art models. For instance, Zhang et al. (2021) demonstrate that modern architectures can perfectly fit datasets with randomly assigned labels, thereby achieving 100 % training accuracy in the absence of any learnable structure. In such cases, high accuracy is attainable only through memorization of individual training samples, underscoring that contemporary artificial neural networks (ANNs) can encode sample-specific and task-irrelevant information to fit each training sample individually.  \nThis ability to memorize arbitrary labels is directly connected to the model complexity. In particular, training with SGD on random labels empirically approximates Rademacher complexity, which plays a central role in deriving generalization bounds within the PAC-learning framework.  \nThe primary objective of this work is to assess the accuracy of predicting random labels as a practical metric of memorization. Although direct training on random labels reveals a model’s ability to memorize, this procedure does not intrinsically inform how memorization interacts with generalization in real-world tasks and does not allow memorization mitigation. To bridge this gap, we propose a hybrid approach: we augment the network with an additional Random Label Prediction Head (RLPhead), attached to the feature extractor (i.e., all layers except the final classification layer) in parallel to the original task head, which remains unchanged. This design enables simult","cbCaia9eslzaKo4r","https://ap.wps.com/l/cbCaia9eslzaKo4r","pdf",1475469,4,1,24,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What is the core idea behind Random Label Prediction (RLP) heads?\",\"answer\":\"RLP-heads augment training by predicting auxiliary random labels from intermediate representations, enabling analysis of memorization at different network depths during normal training.\"},{\"question\":\"How does RLP-head accuracy connect to memorization and model capacity?\",\"answer\":\"RLP-head performance is interpreted as an empirical estimate of Rademacher complexity, providing a practical measure of both sample-level memorization and overall model capacity.\"},{\"question\":\"How is memorization mitigated in the proposed approach?\",\"answer\":\"A new regularization technique constrains memorization by penalizing the RLP-head output performance during training, demonstrably reducing memorization.\"}]",1784209858,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"random-label-prediction-heads-for-studying-memorization-in-deep-neural-networks","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/random-label-prediction-heads-for-studying-memorization-in-deep-neural-networks/86257/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea behind Random Label Prediction (RLP) heads?","Question",{"text":75,"@type":76},"RLP-heads augment training by predicting auxiliary random labels from intermediate representations, enabling analysis of memorization at different network depths during normal training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RLP-head accuracy connect to memorization and model capacity?",{"text":80,"@type":76},"RLP-head performance is interpreted as an empirical estimate of Rademacher complexity, providing a practical measure of both sample-level memorization and overall model capacity.",{"name":82,"@type":73,"acceptedAnswer":83},"How is memorization mitigated in the proposed approach?",{"text":84,"@type":76},"A new regularization technique constrains memorization by penalizing the RLP-head output performance during training, demonstrably reducing memorization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]