[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126359-en":3,"doc-seo-126359-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126359,962085564381,"Clementine","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Machine Learning-based Reentry Guidance for the ReFEx Mission - Analysis and Comparison of Reinforcement Learning and Genetic Programming","This thesis develops and evaluates machine-learning guidance laws for safe atmospheric reentry of the ReFEx winged reusable launch vehicle. Reinforcement Learning and Genetic Programming are applied to generate real-time updates of angle of attack and bank angle commands once a dynamic pressure threshold is reached. Models trained in simulation are validated using a high-fidelity 6-DOF environment developed by DLR, and the resulting performance is compared across the two approaches to support selection of an algorithm suitable for mission certification.","Institute of Space Systems  \nExecutive Summary of the Thesis  \nMachine Learning-based Reentry Guidance for the ReFEx Mission: Analysis and Comparison of Reinforcement Learning and Genetic Programming  \nLaurea Magistrale in Space Engineering-Ingegneria Spaziale  \nAuthor: Riccardo Cadamuro  \nAdvisor: Prof. Francesco Topputo  \nCo-advisor: Dr. Francesco Marchetti  \nAcademic year: 2024-2025  \n1. Introduction and Objectives  \nThe reusability of launch vehicles is attracting increasing interest within the space industry, primarily due to its potential to significantly reduce mission costs. However, guiding a vehicle safely through atmospheric reentry from space remains a highly complex Guidance and Control (G&C) challenge, largely due to the uncertain and rapidly changing environmental conditions. As a result, the development of efficient and reliable G&C algorithms has become essential. Machine Learning (ML) techniques are gaining considerable attention, as they enable the offline training of models in simulated environments, which once deployed in real systems, can generate guidance commands in real-time with minimal computational overhead.  \nThe objective of this thesis is to apply Reinforcement Learning (RL) and Genetic Programming (GP) to the design of the reentry guidance law of the REusability Flight Experiment (ReFEx) mission: a winged Reusable Launch Vehicle (RLV) designed according to the Vertical Takeoff Horizontal Landing (VTHL) paradigm to follow a realistic trajectory for a general winged RLV. The ReFEx baseline guidance algorithm  \ncomputes the initial reference commands for angle of attack (α) and bank angle (µ) onboard prior to the start of the reentry phase [3] . Once reentry begins (i.e. a given dynamic pressure threshold is met), these reference commands are iteratively updated in real time by applying corrections δα and δµ until the End of Experiment (EoE) is reached [4] . This work aims to: 1) use RL and GP to train a model capable of generating online the guidance command update during the atmospheric reentry phase; 2) assess the performances of the trained ML models using the high-fidelity 6-DOF simulator, internally developed by DLR, used to certify the state-of-the-art algorithm designed for the real mission; and 3) produce a comparative study between RL and GP.  \n2. Reinforcement Learning  \nRL is a technique used to train an agent that learns to compute optimal actions through interaction with an environment, aiming to maximize a numerical reward signal. The problem is framed as a Markov Decision Process (MDP), where the decision maker (agent) observes the system’s state as it evolves over time and selects  \nExecutive summary Riccardo Cadamuro  \nactions from a feasible set at each time step [6] . The function mapping states into feasible actions is called policy and it is learnt during the training process with the objective of maximizing the so-called cumulative discounted reward in Equation 1 .  \nGt = Rt+1 + γRt+2 + γ2 Rt+3 + ... = P γkRi+t+1  \n(1)  \nThe discount rate γ ∈ [0 , 1] controls the relative importance of future rewards (γ ≈ 1) versus immediate ones (γ ≈ 0) . The training is performed over episodic tasks, where each episode corresponds to a full trajectory from Entry Interface (EI) to EoE. Once the episode is complete, the initial conditions are reset and a new one is begun.  \nThis work uses the Proximal Policy Optimization (PPO) algorithm to train a Multilayer Perceptron (MLP). The PPO algorithm [5] is chosen because it is designed to deal with a continuous action space and to perform small and controlled policy update steps which strongly stabilize the training performances. It can learn the optimal behavior by interacting with the environment without having knowledge of its model (model free method) and it simultaneously trains two networks: a first one to compute the action anda second one to evaluate the action produced by the first (Actor-Critic method) . On the other hand, the MLP has bee","cbCaid4XpGXpNyuq","https://ap.wps.com/l/cbCaid4XpGXpNyuq","pdf",629572,1,6,"English","en",105,"# Introduction and Objectives\n## Reinforcement Learning\n## Genetic Programming","[{\"question\":\"What is the main objective of the thesis?\",\"answer\":\"To design a reentry guidance law for the ReFEx mission using reinforcement learning and genetic programming, and to compare their performance for real-time command updates.\"},{\"question\":\"How does the baseline ReFEx guidance algorithm work?\",\"answer\":\"It onboard-computes initial reference commands for angle of attack and bank angle before reentry, then iteratively updates these commands in real time with correction terms until the end of the experiment.\"},{\"question\":\"Why are reinforcement learning and genetic programming compared in this work?\",\"answer\":\"Reinforcement learning trains a policy to produce continuous guidance actions, while genetic programming evolves symbolic program trees; comparing them highlights differences in training behavior, model structure, and guidance performance under the same high-fidelity validation setup.\"}]","Machine Learning-based Reentry Guidance for the ReFEx Mission - Analysis and Comparison of Reinforcement Learning and Genetic Programming | PDF",1785904656,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"machine-learning-based-reentry-guidance-for-the-refex-mission-analysis-and-comparison-of-reinforcement-learning-and-genetic-programming","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-based-reentry-guidance-for-the-refex-mission-analysis-and-comparison-of-reinforcement-learning-and-genetic-programming/126359/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main objective of the thesis?","Question",{"text":76,"@type":77},"To design a reentry guidance law for the ReFEx mission using reinforcement learning and genetic programming, and to compare their performance for real-time command updates.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the baseline ReFEx guidance algorithm work?",{"text":81,"@type":77},"It onboard-computes initial reference commands for angle of attack and bank angle before reentry, then iteratively updates these commands in real time with correction terms until the end of the experiment.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are reinforcement learning and genetic programming compared in this work?",{"text":85,"@type":77},"Reinforcement learning trains a policy to produce continuous guidance actions, while genetic programming evolves symbolic program trees; comparing them highlights differences in training behavior, model structure, and guidance performance under the same high-fidelity validation setup.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]