[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124912-en":3,"doc-seo-124912-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124912,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","How to choose the right transfer learning protocol - A qualitative analysis in a controlled set-up","Transfer learning enables training with limited data, making it essential for real-world settings where data acquisition is slow and costly. Standard protocols often freeze feature-extractor layers from a pretrained source task and adapt only task-specific readout layers, relying on assumptions about qualitative feature-map similarity and the expressiveness of representations from the deepest hidden layers. This study shows these assumptions do not always hold, and larger gains can arise from transferring smaller network portions. Experiments in controlled conditions reveal that optimal transfer depth depends on training-data availability and source-target similarity, and that transfer-first-layers strategies are often advantageous, alongside a method to select promising source tasks.","How to choose the right transfer learning protocol? A qualitative analysis in a controlled set-up  \nFederica Gerace  \nScuola Internazionale Superiore di Studi Avanzati (SISSA) Department of Mathematics, University of Bologna  \nDiego Doimo  \nArea Science Park  \nStefano Sarao Mannelli  \nGatsby Computational Neuroscience Unit & Sainsbury Wellcome Centre University College London  \nLuca Saglietti  \nDepartment of Computing Sciences Bocconi University  \nAlessandro Laio  \nScuola Internazionale Superiore di Studi Avanzati (SISSA)  \nThe Abdus Salam International Centre for Theoretical Physics (ICTP)  \n[federica.gerace@unibo.it](federica.gerace@unibo.it)  \n[diego. doimo@areasciencepark.it](diego. doimo@areasciencepark.it)  \n[s. saraomannelli@ucl. ac.uk](s. saraomannelli@ucl. ac.uk)  \n[luca.saglietti@unibocconi.it](luca.saglietti@unibocconi.it)  \n[laio@sissa.it](laio@sissa.it)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= XWQgXLYwv2](https: // openreview. net/ forum? id= XWQgXLYwv2)  \nAbstract  \nTransfer learning is a powerful technique that enables model training with limited amounts of data, making it crucial in many data-scarce real-world applications. Typically, transfer learning protocols require first to transfer all the feature-extractor layers of a network pretrained on a data-rich source task, and then to adapt only the task-specific readout layers to a data-poor target task. This workflow is based on two main assumptions: first, the feature maps of the pre-trained model are qualitatively similar to the ones that would have been learned with enough data on the target task; second, the source representations of the last hidden layers are always the most expressive. In this work, we demonstrate that this is not always the case and that the largest performance gain may be achieved when smaller portions of the pre-trained network are transferred. In particular, we perform a set of numerical experiments in a controlled setting, showing how the optimal transfer depth depends non-trivially on the amount of available training data and on the degree of sourcetarget task similarity, and it is often convenient to transfer only the first layers. We then propose a strategy to detect the most promising source task among the available candidates.  \nThis approach compares the internal representations of a network trained entirely from scratch on the target task with those of the networks pre-trained on the potential source tasks.  \n1 Introduction  \nMachine learning models show a remarkable capacity to extrapolate rules and predict the behavior of complex systems. Still, this ability often comes at the cost of training with large amounts of data. In various domains – such as in medical applications – data collection is a slow and costly process. This goes at odds with the typical vast variability across data, which inevitably requires the usage of huge datasets to achieve satisfying  \ngeneralization performance on a desired task Beam & Kohane (2018); Rajkomar et al. (2019) . Data efficiency is thus a necessary condition.  \nTransfer learning emerged as an efficacious mitigation strategy to this problem Thrun & Pratt (2012); Shinet al. (2016); Raghu et al. (2019) . Transferring the already meaningful representations learned on the source task to a network that has to solve a given target task allows it to work with dramatically smaller dataset sizes while keeping a comparable level of accuracy. With little data at disposal, one of the major risks is to incur overfitting Geirhos et al. (2020) . However, transferring the layers directly from another network allows, in principle, to filter out irrelevant information present in the input of both the source and the target datasets and work with a representation of the data of reduced efficacious dimension. This allows a significant increase in data efficiency.  \nUnfortunately, deep neural networks tend to be sensitive even to tiny distribution shifts Gama et al. (2014) . Therefore, i","cbCaicS3uPHBWzyB","https://ap.wps.com/l/cbCaicS3uPHBWzyB","pdf",1279232,1,18,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What typical assumptions do standard transfer learning protocols rely on?\",\"answer\":\"They assume the pretrained model’s feature maps are qualitatively similar to target-task features, and that representations in the last hidden layers are always the most expressive.\"},{\"question\":\"Why might transferring all layers not lead to the best performance?\",\"answer\":\"Because deep models are sensitive to distribution shifts, and the study shows the largest performance gains may come from transferring smaller portions of the pretrained network rather than all the way up to the last hidden representation.\"},{\"question\":\"How does the optimal transfer depth relate to training data and task similarity?\",\"answer\":\"In controlled numerical experiments, the optimal transfer depth depends non-trivially on how much target training data is available and on the degree of similarity between the source and target tasks.\"}]","How to choose the right transfer learning protocol - A qualitative analysis in a controlled set-up | PDF",1785895355,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"how-to-choose-the-right-transfer-learning-protocol-a-qualitative-analysis-in-a-controlled-set-up","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/how-to-choose-the-right-transfer-learning-protocol-a-qualitative-analysis-in-a-controlled-set-up/124912/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What typical assumptions do standard transfer learning protocols rely on?","Question",{"text":75,"@type":76},"They assume the pretrained model’s feature maps are qualitatively similar to target-task features, and that representations in the last hidden layers are always the most expressive.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why might transferring all layers not lead to the best performance?",{"text":80,"@type":76},"Because deep models are sensitive to distribution shifts, and the study shows the largest performance gains may come from transferring smaller portions of the pretrained network rather than all the way up to the last hidden representation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the optimal transfer depth relate to training data and task similarity?",{"text":84,"@type":76},"In controlled numerical experiments, the optimal transfer depth depends non-trivially on how much target training data is available and on the degree of similarity between the source and target tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]