[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84584-en":3,"doc-seo-84584-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84584,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Loss Smoothing for Stable Adaptation Under Distribution Shift","Neural networks are adapted under distribution shift in fine-tuning and reinforcement learning, where standard methods abruptly switch to the target objective, distorting representations that may still help the new task. The work proposes loss smoothing, which replaces the discontinuous transition with a convex interpolation between source and target objectives at adaptation start, then anneals to the target loss. Experiments across supervised shifts, vision adaptation, offline-to-online and online RL, and language-model finetuning show consistent performance gains.","arXiv :2607 .00634v 1 [ cs .LG] 1 Jul 2026  \nLoss Smoothing for Stable Adaptation Under Distribution Shift  \nDarshan Patil 1 ,2 ,3 Ekaterina Lobacheva 1 ,2 ,3 Razvan Pascanu2 Sarath Chandar 1 ,2 ,4 ,5  \n1 Chandar Research Lab 2Mila – Quebec AI Institute 3Université de Montréal  \n4Polytechnique Montréal 5 Canada CIFAR AI Chair  \nCorrespondence: [darshan.patil@mila.quebec](darshan.patil@mila.quebec)  \nAbstract  \nIn settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation methods typically optimize the target objective directly, inducing an abrupt change from the source training objective. This abrupt transition can distort learned representations, including features that may still be useful for the new task. We investigate whether a more gradual transition can improve adaptation. We propose loss smoothing, a simple approach that interpolates between the source and target training objectives at the start of adaptation. This smooth transition helps to preserve useful features from the source distribution while still enabling the model to specialize to the target distribution. Across controlled supervised shifts, pretrained vision adaptation, offline-to-online and online reinforcement learning, and language model finetuning, we find that loss smoothing consistently improves performance, suggesting that smoother objective transitions are a broadly useful tool for model adaptation.  \n1 Introduction  \nModern neural networks are rarely trained once and then left unchanged. A pretrained language model is adapted to instruction-following data [Ouyang et al., 2022, Wang et al., 2023]; a classifier trained on one task is updated for the next [Kirkpatrick et al., 2017, Sodhani et al., 2022]; a policy trained from offline data is fine-tuned through environment interaction [Nair et al., 2021, Nakamoto et al., 2023, Shin et al., 2025] . In each case, adaptation begins from a model that already contains useful features, but then exposes that model to a new input data distribution, a new learning signal, or even a new objective.  \nThe adaptation problem is therefore not simply to optimize the new objective, but to do so while making effective use of the representations already present in the model. This perspective changes the role of stability in adaptation. Stability is often treated as the opposite of adaptation: a model that stays too close to its source solution may fail to fit the new task, but the goal is not simply to move as much as possible. Many shifts contain both task-shared components, such as reusable features, value estimates, behaviors, or representations, and task-inconsistent components that conflict with the new objective. Successful adaptation should preserve the former while modifying the latter. For example, an offline RL policy may contain useful exploratory behavior even when the offline objective is too conservative for online improvement; a pretrained language model may contain broad general-purpose capabilities even when instruction tuning pushes it toward a narrow response distribution. Prior work has observed related brittleness in feature-distorting fine-tuning [Kumar et al., 2022], offline-to-online RL [Nakamoto et al., 2023, Shin et al., 2025], and overtrained language-model fine-tuning [Springer et al., 2025] .  \nWe argue brittle adaptation can arise not only because the target objective is difficult, but because the transition into it is poorly chosen. In standard fine-tuning, once the new phase begins, the source  \nPreprint.  \nobjective is discarded and the model is optimized directly on the target objective. This hard switch can create large, poorly directed updates exactly when the model has the least information about which learned representations should be preserved.  \nWe propose loss smoothing, a simple method for replacing this discontinuous transition with a smoother objective transition. At the start of an adaptation phase,","cbCaibcLKeD8pE0J","https://ap.wps.com/l/cbCaibcLKeD8pE0J","pdf",1263259,1,26,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the paper target in adaptation under distribution shift?\",\"answer\":\"It targets the abrupt switch to the target objective in standard adaptation, which can distort learned representations and hinder stable reuse of useful features.\"},{\"question\":\"How does loss smoothing work during the start of adaptation?\",\"answer\":\"At the beginning of adaptation, it optimizes a convex interpolation between the source and target objectives, then anneals the interpolation coefficient until standard target training is recovered.\"},{\"question\":\"Where does loss smoothing improve performance according to the experiments?\",\"answer\":\"Loss smoothing improves performance across controlled supervised shifts, pretrained vision adaptation, offline-to-online and online reinforcement learning, and language model finetuning.\"}]",1784196934,66,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"loss-smoothing-for-stable-adaptation-under-distribution-shift","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/loss-smoothing-for-stable-adaptation-under-distribution-shift/84584/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper target in adaptation under distribution shift?","Question",{"text":74,"@type":75},"It targets the abrupt switch to the target objective in standard adaptation, which can distort learned representations and hinder stable reuse of useful features.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does loss smoothing work during the start of adaptation?",{"text":79,"@type":75},"At the beginning of adaptation, it optimizes a convex interpolation between the source and target objectives, then anneals the interpolation coefficient until standard target training is recovered.",{"name":81,"@type":72,"acceptedAnswer":82},"Where does loss smoothing improve performance according to the experiments?",{"text":83,"@type":75},"Loss smoothing improves performance across controlled supervised shifts, pretrained vision adaptation, offline-to-online and online reinforcement learning, and language model finetuning.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]