[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122268-en":3,"doc-seo-122268-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122268,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Update Your Transformer to the Latest Release - Re-Basin of Task Vectors","Foundation models underpin many specialized systems obtained via fine-tuning, yet updated pretrained checkpoints can quickly make existing fine-tuned models obsolete. The work studies training-free transfer of fine-tuning to a new model release by re-basing task vectors onto the new checkpoint. It adapts model re-basin for Transformer architectures using a two-level spectral approach that permutes attention heads and tunes parameters in selected head pairs. Experiments on visual and textual tasks show seamless transfer without any training steps or datapoints, with code released publicly.","This is the peer reviewd version of the followng article:  \nUpdate Your Transformer to the Latest Release: Re-Basin of Task Vectors / Rinaldi, Filippo; Capitani, Giacomo; Bonicelli, Lorenzo; Crisostomi, Donato; Bolelli, Federico; Rodolà, Emanuele; Ficarra, Elisa; Calderara, Simone; Porrello, Angelo. - (2025) . (Intervento presentato al convegno International Conference on Machine Learning tenutosi a Vancouver, Canada nel Jul 13-19) .  \nTerms of use:  \nThe terms and conditions for the reuse of this version of the manuscript are specified in the publishing policy. For all terms of use and more information see the publisher's website.  \n12/06/2025 11:04  \n(Article begins on next page)  \nUpdate Your Transformer to the Latest Release: Re-Basin of Task Vectors  \nFilippo Rinaldi * 1 Giacomo Capitani * 1 Lorenzo Bonicelli 1 Donato Crisostomi 2 Federico Bolelli 1 Elisa Ficarra 1 Emanuele Rodol2 Simone Calderara 1 Angelo Porrello 1  \nAbstract  \nFoundation models serve as the backbone for numerous specialized models developed through fine-tuning. However, when the underlying pretrained model is updated or retrained (e.g., on larger and more curated datasets), the fine-tuned model becomes obsolete, losing its utility and requiring retraining. This raises the question: is it possible to transfer fine-tuning to a new release of the model? In this work, we investigate how to transfer fine-tuning to a new checkpoint without having to re-train, in a data-free manner. Todo so, we draw principles from model re-basin and provide a recipe based on weight permutations to re-base the modifications made to the original base model, often called task vector. In particular, our approach tailors model re-basin for Transformer models, taking into account the challenges of residual connections and multi-head attention layers. Specifically, we propose a twolevel method rooted in spectral theory, initially permuting the attention heads and subsequently adjusting parameters within select pairs of heads.  \nThrough extensive experiments on visual and textual tasks, we achieve the seamless transfer of fine-tuned knowledge to new pre-trained backbones without relying on a single training step or datapoint. Code is available at [https://](https://)[ ](https://)[github.com/aimagelab/TransFusion](github.com/aimagelab/TransFusion).  \n1. Introduction  \nRecently, there has been a notable shift among researchers and practitioners towards fine-tuning pre-trained models, rather than building them from scratch. This method leverages backbones developed on large-scale datasets, consider-  \n*Equal contribution 1AImageLab, University of Modena and Reggio Emilia, Italy. 2 Sapienza, University of Rome, Italy. Correspondence to: Filippo Rinaldi \u003C[filippo.rinaldi@unimore.it](filippo.rinaldi@unimore.it) >.  \nProceedings of the 42 st International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nFigure 1: Transporting task vector τ from a fine-tuned base model θftA = θA + τ to a new release θB .  \nably decreasing the amount of data and training time needed to tailor models for specific downstream tasks. For this reason, pre-trained backbones such as OpenAI’s CLIP (Radford et al., 2021) are being extensively utilized as base foundation models. As a result, the corresponding fine-tuned versions play a crucial role in numerous real-world applications like medical imaging (Lu et al., 2024) and satellite image analysis (Mall et al., 2024) .  \nHowever, while these pre-trained backbones are widely adopted, their evolution poses new challenges, with tech companies and academic institutions frequently releasing updated checkpoints. Often, these updates do not modify the underlying architecture but simply consist of new weights, trained on increasingly large datasets compared to their predecessors (Ilharco et al., 2021) . Moreover, the additional training data may be more curated or specifically tailored to specialized domains, boosting ","cbCaiao1rspnLcnj","https://ap.wps.com/l/cbCaiao1rspnLcnj","pdf",708598,1,16,"English","en",105,"# Introduction\n## Motivation: fine-tuning becomes obsolete after checkpoint updates\n## Goal: data-free training-free transfer of fine-tuning\n## Task vector transport via model re-basin for Transformers\n## Proposed two-level spectral method\n## Experimental evaluation and results","[{\"question\":\"What problem does the paper address when new pretrained checkpoints are released?\",\"answer\":\"When the underlying pretrained model is updated or retrained, previously fine-tuned models can become obsolete, requiring costly retraining.\"},{\"question\":\"How does the proposed method transfer fine-tuning to a new model release without training?\",\"answer\":\"It transports the task vector from the original base model to a suitable basin of the new release using weight permutations, re-basing the modifications without training steps or datapoints.\"},{\"question\":\"What are the main components of the Transformer-specific re-basin approach?\",\"answer\":\"The method uses a two-level strategy based on spectral theory: first permuting attention heads, then adjusting parameters within selected pairs of heads to improve compatibility and low-loss transfer.\"}]","Update Your Transformer to the Latest Release - Re-Basin of Task Vectors | PDF",1785809749,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"update-your-transformer-to-the-latest-release-re-basin-of-task-vectors","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/update-your-transformer-to-the-latest-release-re-basin-of-task-vectors/122268/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address when new pretrained checkpoints are released?","Question",{"text":75,"@type":76},"When the underlying pretrained model is updated or retrained, previously fine-tuned models can become obsolete, requiring costly retraining.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method transfer fine-tuning to a new model release without training?",{"text":80,"@type":76},"It transports the task vector from the original base model to a suitable basin of the new release using weight permutations, re-basing the modifications without training steps or datapoints.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main components of the Transformer-specific re-basin approach?",{"text":84,"@type":76},"The method uses a two-level strategy based on spectral theory: first permuting attention heads, then adjusting parameters within selected pairs of heads to improve compatibility and low-loss transfer.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]