[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83227-en":3,"doc-seo-83227-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83227,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Trajectory-selection RL MPC for Greenhouse Fruit-production Control","Greenhouse fruit-production control targets maximum economic return (fruit revenue minus operating costs) while satisfying system constraints under uncertain external weather disturbances. The work addresses the trade-off between MPC’s constraint handling with short horizons and RL’s ability to incorporate long-term economic objectives but weaker constraint enforcement and robustness to unseen weather. A trajectory-selection RL–MPC framework blends an RL rollout with short-horizon nonlinear MPC via terminal region and cost, selecting the better objective candidate for control.","Graphical Abstract  \nImproving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control  \nB. van Laatum, S. Msaad, E. J. van Henten, R.D. McAllister, S. Boersma  \ndt:t+N-1  \nTrajectory-selection RL-MPC  \nWeather  \n~~ ~~ Predicted trajectory (x★, u★)  \n~~ ~~ RL rollout trajectory (x, u)  \nDiscrete-time model 􀥿f (xt+N)  \nut  \nπRL xt+N  \n[math .OC] 8 Jul 2026  \narXiv :2607 .07365v1  \nHighlights  \nImproving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control  \nB. van Laatum, S. Msaad, E. J. van Henten, R.D. McAllister, S. Boersma  \n• Trajectory-selection RL–MPC improves short-horizon MPC for greenhouse production control  \n• Applied to GreenLight, a large-scale nonlinear greenhouse tomato production model  \n• RL policy rollouts incorporate long-term economic information into short-horizon MPC  \n• Improved closed-loop performance over RL policy on unseen weather trajectories  \n• The source code is publicly available at [https://github.com/BartvLaatum/GL-Gym-MPC](https://github.com/BartvLaatum/GL-Gym-MPC)  \nImproving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control B. van Laatuma,∗ , S. Msaadb , E. J. van Hentena , R.D. McAllisterb and S. Boersmaa,c  \na Agricultural Biosystems Engineering, Wageningen University & Research, Droevendaalsesteeg 4, Wageningen, 6708 PB, The Netherlands b Delft Center for Systems and Control (DCSC), Delft University of Technology, Mekelweg 5, Delft, 2628 CD, The Netherlands c Biometris, Wageningen Research, Droevendaalsesteeg 4, Wageningen, 6708 PB, The Netherlands  \n\n| ARTICLE INFO\u003Cbr>Keywords:\u003Cbr>Model predictive control Reinforcement learning Greenhouse fruit production control Climate control\u003Cbr>Finite-horizon optimization | AB STRACT\u003Cbr>Greenhouse fruit-production control aims to maximize the economic performance (fruit revenue minus operating costs) while operating within system constraints under external weather disturbances. Control methods need to balance the delayed economic benefit of fruit yield with current operating costs. For such problems, model predictive control (MPC) can explicitly handle system constraints under future weather disturbances, but can become computationally demanding when using sufficiently long prediction horizons for (relatively large) nonlinear greenhouse fruit production models. In contrast, reinforcement learning (RL) can learn control policies offline while considering longer-term economic performance, but struggles to enforce system constraints, and performance may degrade under unseen weather trajectories. This work proposes trajectory-selection RL–MPC, a framework that incorporates longer-term economic information of fruit yield into a short-horizon MPC optimization problem. The framework uses an RL rollout trajectory to define a terminal region constraint and terminal cost. Next, a nonlinear MPC solves a short-horizon optimization problem with these terminal ingredients to find a local optimum. Finally, the framework selects and executes the first input from the trajectory with the better objective value, either from the MPC-predicted or the RL rollout trajectory. The method is applied to GreenLight, a large-scale greenhouse tomato production model that exhibits stiff dynamics. The simulation results show that trajectory-selection RL–MPC with a one-hour prediction horizon matches the closed-loop performance of a high-performing guiding policy while significantly improving over standalone MPC with the same horizon. On 25 weather trajectories held out during policy training, trajectory-selection RL–MPC achieved 54% higher closed-loop performance than standalone RL and 80% higher performance than MPC with the same horizon length. |\n| --- | --- |\n| 1. Introduction\u003Cbr>Changing climate conditions and more extreme weather events threaten the stable production of fruits and veg","cbCaimhPJ95JVgn2","https://ap.wps.com/l/cbCaimhPJ95JVgn2","pdf",2005748,3,1,27,"English","en",105,"# Introduction\n## Trajectory-selection RL–MPC framework\n## Application to the GreenLight greenhouse tomato model\n## Simulation results and performance comparison","[{\"question\":\"What is the main objective of greenhouse fruit-production control in this work?\",\"answer\":\"Maximize economic performance defined as fruit revenue minus operating costs while keeping the system within constraints under external weather disturbances.\"},{\"question\":\"Why does the paper combine reinforcement learning with short-horizon model predictive control (MPC)?\",\"answer\":\"MPC can handle constraints under future weather but becomes computationally heavy with long horizons, while RL can optimize long-term economics offline but struggles to enforce constraints and may degrade on unseen weather. The proposed method injects long-term economic information into a short-horizon MPC problem.\"},{\"question\":\"How does trajectory-selection RL–MPC choose actions at each step?\",\"answer\":\"It generates an RL rollout trajectory that defines a terminal region constraint and terminal cost, runs short-horizon nonlinear MPC to compute a local optimum trajectory, then selects and executes the first input from whichever candidate trajectory yields the better objective value.\"}]",1784186072,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"trajectory-selection-rl-mpc-for-greenhouse-fruit-production-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/trajectory-selection-rl-mpc-for-greenhouse-fruit-production-control/83227/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main objective of greenhouse fruit-production control in this work?","Question",{"text":75,"@type":76},"Maximize economic performance defined as fruit revenue minus operating costs while keeping the system within constraints under external weather disturbances.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the paper combine reinforcement learning with short-horizon model predictive control (MPC)?",{"text":80,"@type":76},"MPC can handle constraints under future weather but becomes computationally heavy with long horizons, while RL can optimize long-term economics offline but struggles to enforce constraints and may degrade on unseen weather. The proposed method injects long-term economic information into a short-horizon MPC problem.",{"name":82,"@type":73,"acceptedAnswer":83},"How does trajectory-selection RL–MPC choose actions at each step?",{"text":84,"@type":76},"It generates an RL rollout trajectory that defines a terminal region constraint and terminal cost, runs short-horizon nonlinear MPC to compute a local optimum trajectory, then selects and executes the first input from whichever candidate trajectory yields the better objective value.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]