[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85055-en":3,"doc-seo-85055-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85055,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Understanding Layer Patching in Model Size Interpolation","Zero-shot model size interpolation constructs intermediate-size language models by combining existing models without extra training. Building on boomerang distillation, where student layers are replaced with contiguous teacher-layer blocks to smoothly interpolate capacity and performance, the study targets a key open problem: selecting which student layers to patch. The work formulates the selection as an optimization problem, proves an equivalent shortest-path view on an acyclic graph, and shows patching order strongly shapes interpolation across model families. Sequential and KL-based greedy methods often achieve near-optimal results.","arXiv :2607 .08 170v 1 [ cs .LG] 9 Jul 2026  \nUnderstanding Layer Patching in Model Size Interpolation  \nSara Kangaslahti1 , Jonathan Geuter1  Nihal V. Nayak1,2  Marco Fumero3   \nFrancesco Locatello3 , David Alvarez-Melis 1,2  \n1Harvard University 2 Kempner Institute 3IST Austria  \n[sarakangaslahti@g.harvard.edu](sarakangaslahti@g.harvard.edu)  \nAbstract  \nZero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distillation [Kangaslahti et al., 2026] shows that a student language model distilled from a larger teacher can be expanded by iteratively patching its layers, replacing student layers with contiguous blocks of teacher layers to obtain models whose size and performance interpolate between the student and the teacher. Selecting which layers to patch for a given intermediate model size is a key design choice of this procedure, yet it has remained largely underexplored. In this work, we provide the first systematic study of student-layerselection for model size interpolation. We cast finding the optimal layer subset for each model size as an optimization problem and prove it can be viewed as a shortest-path problem in a certain acyclic graph. In experiments, we show that patching strongly shapes interpolation behavior, with effects that vary substantially across model families. We find that simple sequential strategies—patching either from the first layer to the last or from the last to the first—often achieve surprisingly strong performance in practice. We further introduce KLPatch, a greedy patching algorithm based on KL divergence, which often improves over last-to-first patching and approximately solves the optimization problem. Together, our results provide a principled understanding of how layer patching affects model size interpolation and offer practical guidance for constructing near-optimal interpolated models.  \n1 Introduction  \nCreating LLM families with different model sizes is one of the most practical ways to adapt to users’varying compute constraints [Huyen, 2022, Khandelwal et al., 2025, Team et al., 2026] . However, pre-training LLMs of different sizes requires separate training runs, which often results in the release of only a small number of coarse-grained, fixed-size LLMs [Grattafiori et al., 2024, Yang et al., 2025, Team, 2026] . Recent work on boomerang distillation [Kangaslahti et al., 2026] has explored zero-shot model-size interpolation, enabling the creation of LLM families with fine-grained sizes. Boomerang distillation is a procedure in knowledge distillation in which layers of a student LLM can be patched with contiguous blocks of teacher layers to create intermediate-size models whose performance smoothly interpolates between that of the student and the teacher LLM, without requiring additional training. Boomerang distillation involves several critical choices, such as the student initialization, training token budget, alignment losses, and student patching. Among these, the role of student patching, a key user-facing decision with a combinatorial design space, remains poorly understood.  \nPrior work on boomerang distillation provides limited guidance on how to patch student models for optimal interpolation performance. Kangaslahti et al. [2026] consider only two patching strategies:  \n∗Equal contribution.  \nPreprint.  \npatching from the first layer to the last and patching from the last layer to the first. Although they show that some LLMs prefer one order over the other, the broader design space of patching remains largely unexplored. Moreover, in Section 5, we show that naively increasing the size of an interpolated model by arbitrarily patching the student can degrade performance, contradicting the conventional view that larger models perform better [Sutton, 2019, Kaplan et al., 2020, Hoffmann et al., 2022] . These observations motivate a deeper understanding of patching t","cbCaioNw8B91ERyI","https://ap.wps.com/l/cbCaioNw8B91ERyI","pdf",1732965,1,31,"English","en",105,"# Introduction\n# Layer Patching as Optimization\n# Shortest-Path Formulation\n# Experiments and Results\n# KLPatch Algorithm","[{\"question\":\"What is zero-shot model size interpolation in this paper?\",\"answer\":\"It creates intermediate target-size language models by combining an existing student and teacher without additional training, aiming for smooth performance interpolation across sizes.\"},{\"question\":\"Why is layer patch selection important for interpolation quality?\",\"answer\":\"The paper shows that which student layers to patch is a key design choice: patching order can substantially change interpolation behavior and performance, sometimes even degrading results if done naively.\"},{\"question\":\"How does the paper solve the problem of choosing layer subsets for each target size?\",\"answer\":\"It frames the layer selection as an optimization problem, proves it can be interpreted as a shortest-path problem in a directed acyclic graph, and introduces KLPatch, a greedy KL-divergence-based algorithm that approximates the solution with lower computational cost.\"}]",1784200679,78,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"understanding-layer-patching-in-model-size-interpolation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/understanding-layer-patching-in-model-size-interpolation/85055/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is zero-shot model size interpolation in this paper?","Question",{"text":75,"@type":76},"It creates intermediate target-size language models by combining an existing student and teacher without additional training, aiming for smooth performance interpolation across sizes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is layer patch selection important for interpolation quality?",{"text":80,"@type":76},"The paper shows that which student layers to patch is a key design choice: patching order can substantially change interpolation behavior and performance, sometimes even degrading results if done naively.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper solve the problem of choosing layer subsets for each target size?",{"text":84,"@type":76},"It frames the layer selection as an optimization problem, proves it can be interpreted as a shortest-path problem in a directed acyclic graph, and introduces KLPatch, a greedy KL-divergence-based algorithm that approximates the solution with lower computational cost.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]