[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86006-en":3,"doc-seo-86006-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86006,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Learning to Fine-tune Foundation Models under Resource Limitations","Optimal continual fine-tuning is studied for a pre-trained foundation model deployed on a resource-limited device. At each time slot, sequential data arrive and a controller chooses either to fine-tune—paying compute cost and gaining reward through an application-specific metric—or to discard the batch. The work formulates an online constrained Markov decision process with state defined by model performance, remaining computational budget, and data relevance to past distributions, solved via actor-critic reinforcement learning, with a dynamic programming variant when gains are predictable. Experiments show over 4% higher accuracy with 25% of fine-tuning steps.","Learning to Fine-tune Foundation Models under  \nResource Limitations  \nThomas Tsouparopoulos and Iordanis Koutsopoulos  \nDepartment of Informatics  \nAthens University of Economics and Business  \nAthens, Greece  \narXiv :2607 . 10694v 1 [ cs .LG] 12 Jul 2026  \nAbstract—We study the problem of optimal continual finetuning for a pre-trained Foundation Model deployed at aresource-limited device. At each time slot, a new batch of training data arrives, and the controller is faced with two options: either use the data to fine-tune the model and incur a compute cost, or do not fine-tune the model and discard the data. After the decision, the performance of the current model is measured in terms of an application-specific performance metric such as classification accuracy. Our objective is to learn an optimal policy that determines when to fine-tune the model on a single task (e.g., sentiment analysis), under a finite compute budget. We formulate this online decision-making problem as a constrained Markov Decision Process, where the system state captures three essential aspects: (i) model’s performance,(ii) computational budget, and (iii) data distribution relevance to historic data encountered up to that point. The transition to the next state is stochastic and therefore, we propose a reinforcement learning-based method to solve this problem, namely the actor-critic algorithm. We also consider the special case where the performance of fine-tuning for a given model can be predicted or estimated prior to decision; in this case the problem becomes a Dynamic Programming one. Experiments with a large pre-trained model on a widely-used text classification dataset demonstrate that our method consistently outperforms fine-tuning approaches with the same compute budget by more than 4% in terms of accuracy and achieves 97% of full-parameter fine-tuning accuracy while requiring only 25% of the fine-tuning steps.  \nIndex Terms—Foundation models, Fine-tuning, Continual learning, Reinforcement learning.  \nI. INTRODUCTION  \nFoundation Models (FMs) are large models pre-trained on massive, general-purpose datasets through self-supervised learning, so that they learn general patterns, logical structures, and relationships between concepts. This process creates a flexible general-purpose base model that can then be fine-tuned with smaller datasets for various downstream tasks such as text summarization and code generation, without the need tobe retrained from scratch each time. Fine-tuning (FT) involves updating some or all of the model parameters by performing some training iterations with the new, small dataset.  \nA representative example is a traffic-analysis FM at a 6G base station, initially trained on generic network traces and periodically adapted using locally observed traffic patterns such as new application behaviors or emerging protocol variants. The proposed RL controller learns when to trigger updates based on expected network-level benefit, enabling resourceaware model adaptation that improves traffic classification ac-  \ncuracy while respecting the operational constraints of wireless infrastructure.  \nParameter-efficient fine-tuning (PEFT) methods, e.g., Low Rank Adaptation (LoRA) [1] and adapters are the predominant class of approaches for continually updating large FMs by significantly reducing the number of trainable parameters, allowing for lightweight model updates, while maintaining near full fine-tuning performance at a fraction of the cost. However, the benefit of PEFT methods for a given model changes overtime with data distribution shift, and it depends on the history of FT steps applied to the model’s parameters [2] .  \nWhen a FM resides on a resource-limited device, the problem of FT the base model for a new task obtains an interesting new twist. Training data for FT the model may arrive at the device sequentially, namely in successive data batches. Each batch of data may be exploited for FT the FM, or it maybe discarded. This","cbCaim03hlKJ03ic","https://ap.wps.com/l/cbCaim03hlKJ03ic","pdf",1810428,3,1,6,"English","en",105,"# Introduction\n## Foundation models and fine-tuning\n## Parameter-efficient fine-tuning (PEFT)\n## Continual fine-tuning under resource constraints\n## Online decision-making and challenges","[{\"question\":\"What decision does the controller make at each time slot?\",\"answer\":\"At each time slot, the controller decides whether to fine-tune the foundation model using the newly arrived training batch (incurring compute cost) or to discard the batch.\"},{\"question\":\"How is the problem modeled to compute an optimal fine-tuning policy?\",\"answer\":\"It is formulated as a constrained Markov Decision Process, where the state captures model performance, computational budget, and how relevant the current data distribution is to the history.\"},{\"question\":\"What learning method is proposed for solving the decision process?\",\"answer\":\"The approach uses reinforcement learning with an actor-critic algorithm to handle stochastic transitions and learn the optimal update timing policy.\"}]",1784207730,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-to-fine-tune-foundation-models-under-resource-limitations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learning-to-fine-tune-foundation-models-under-resource-limitations/86006/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What decision does the controller make at each time slot?","Question",{"text":75,"@type":76},"At each time slot, the controller decides whether to fine-tune the foundation model using the newly arrived training batch (incurring compute cost) or to discard the batch.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the problem modeled to compute an optimal fine-tuning policy?",{"text":80,"@type":76},"It is formulated as a constrained Markov Decision Process, where the state captures model performance, computational budget, and how relevant the current data distribution is to the history.",{"name":82,"@type":73,"acceptedAnswer":83},"What learning method is proposed for solving the decision process?",{"text":84,"@type":76},"The approach uses reinforcement learning with an actor-critic algorithm to handle stochastic transitions and learn the optimal update timing policy.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]