[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82408-en":3,"doc-seo-82408-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82408,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","CORAL-AUV: CFD Oriented Reinforcement Learning for Autonomous Underwater Vehicles","Fine-grain control and accurate positioning are essential for autonomous underwater vehicles (AUVs) supporting sampling, maintenance, and survey missions. Traditional controllers are labor intensive and lack robustness to changes in vehicle configuration and environmental conditions. Reinforcement learning can accelerate controller development via domain randomization, but simulation fidelity limits sim-to-real transfer, especially for drag. This work trains surrogate drag models from CFD data to enable efficient zero-shot RL deployment, reducing energy and improving speed and error metrics while strengthening transfer to perturbed parameters.","arXiv :2607 .09557v1 [ cs .RO] 10 Jul 2026  \nCORAL-AUV: CFD Oriented Reinforcement Learning for Autonomous Underwater Vehicles  \nSteven Roche 1 Milo Van Mooy2 Nathan McGuire2 Levi Cai 1 Jonathan P. How3 Yogesh Girdhar2  \n1MIT–WHOI Joint Program  \n2Woods Hole Oceanographic Institution  \n3Massachusetts Institute of Technology  \n{rochesh,[cail](cail}@mit.edu {nmcguire)[}](cail}@mit.edu {nmcguire)[@mit.edu](cail}@mit.edu {nmcguire)[ {](cail}@mit.edu {nmcguire)[nmcguire](cail}@mit.edu {nmcguire) ,[ygirdhar](ygirdhar}@whoi.edu)[}](ygirdhar}@whoi.edu)[@whoi.edu](ygirdhar}@whoi.edu)[ ](ygirdhar}@whoi.edu)[milovm@stanford.edu](milovm@stanford.edu) [jhow@mit.edu](jhow@mit.edu)  \nAbstract: Fine grain control and positioning of autonomous underwater vehicles (AUVs) is critical for sampling, maintenance, and survey applications. Traditional control methods for AUVs are labor intensive and are not robust to changes in the vehicle configuration or environmental conditions. Reinforcement learning (RL) promises rapid controller development while handling a range of deployment parameters via domain randomization (DR) . However, DR is still limited by the capacity of the underlying simulation to model real physics. In particular, drag physics are difficult to model and are a large contributor to sim-to-real gaps.  \nMeanwhile, computational fluid dynamics (CFD) provides high fidelity drag models but is challenging to leverage within reinforcement learning frameworks due to its computational overhead. Thus, in this paper we exploit the idea of training surrogate approximations of CFD models of a given vehicle, enabling fast inference within RL pipelines. We are the first to successfully deploy a zero-shot RL policy on a 6-DOF AUV in which policy training is performed on surrogate drag models (SDMs) trained on CFD data. We find 31% lower energy usage compared to a controller using simplified physics while traversing between waypoints 11% faster with 19% less error. Our SDM based RL controller better predicts zero-shot transfer and is more robust across reward shaping design choices. When using DR to complete a task with perturbed parameters, we find that the CFD policy is the only controller that successfully transfers. The policies are evaluated ina controlled tank environment and in the field providing extensive testing of the policies’ capabilities.  \nKeywords: Sim-to-real transfer, Reinforcement learning for physical robot control, CFD, AUVs  \n1 Introduction  \nAutonomous underwater vehicles (AUVs) have been used for a variety of underwater environmental research interests, from the study of plankton [1], tracking of mobile animal species [2, 3], and surveys of coral reefs [4] . These tasks often require robots to navigate challenging terrain and conditions [5, 6]; for example, a vehicle may need to operate within sub-meter ranges of a fragile coral species to obtain high-resolution visual data [4, 7], or it may need to use multiple sensors to avoid obstacles [8] . Maintaining close proximity to obstacles and reacting quickly to marine life requires fine-tuned control of the dynamics of the vehicles. A mission may require multiple sensor configurations, making controller development extremely time-intensive and challenging due to the changes in the hydrodynamics. For example, Hawkes et al. curtailed one of their AUV experiments prematurely when they added a stern-facing camera to their AUV, causing unintended impacts to their hydrodynamics and thus controller failures [9] . Traditional controllers such as PID may require manual re-tuning, fail to generalize to new tasks, or degrade under changing environmental  \nconditions. Recent works have turned to reinforcement learning (RL) based controllers to address these limitations. Two common approaches to dynamics modeling involve: 1) using high-fidelity computational fluid dynamics (CFD) solvers directly in the RL training loop to obtain an accurate controller [10, 11, 12, 13, 14], or 2) using analy","cbCailVA39JAm75X","https://ap.wps.com/l/cbCailVA39JAm75X","pdf",13436053,3,1,16,"English","en",105,"# Introduction\n## Problem setting and limitations\n## Key idea: surrogate drag models for RL\n# Contributions\n## Zero-shot RL deployment on 6-DOF AUV\n## Effects of reward shaping, domain randomization, and fidelity\n## Sim-to-real evaluation in tank and field","[{\"question\":\"Why do traditional AUV controllers struggle with real missions?\",\"answer\":\"Traditional controllers are labor intensive and often require manual re-tuning, with performance degrading when vehicle configuration or environmental conditions change.\"},{\"question\":\"What limits reinforcement learning with domain randomization for AUVs?\",\"answer\":\"Even with domain randomization, sim-to-real transfer is constrained by the simulation’s ability to model real physics, particularly drag, which is difficult to model accurately.\"},{\"question\":\"How does the approach in CORAL-AUV improve RL deployment?\",\"answer\":\"It trains surrogate drag models from CFD data so reinforcement learning can perform fast inference in RL pipelines, enabling a zero-shot policy trained on surrogate drag while retaining CFD-level drag fidelity.\"}]",1784180170,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"coral-auv-cfd-oriented-reinforcement-learning-for-autonomous-underwater-vehicles","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/coral-auv-cfd-oriented-reinforcement-learning-for-autonomous-underwater-vehicles/82408/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do traditional AUV controllers struggle with real missions?","Question",{"text":75,"@type":76},"Traditional controllers are labor intensive and often require manual re-tuning, with performance degrading when vehicle configuration or environmental conditions change.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limits reinforcement learning with domain randomization for AUVs?",{"text":80,"@type":76},"Even with domain randomization, sim-to-real transfer is constrained by the simulation’s ability to model real physics, particularly drag, which is difficult to model accurately.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the approach in CORAL-AUV improve RL deployment?",{"text":84,"@type":76},"It trains surrogate drag models from CFD data so reinforcement learning can perform fast inference in RL pipelines, enabling a zero-shot policy trained on surrogate drag while retaining CFD-level drag fidelity.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]