[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84535-en":3,"doc-seo-84535-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},84535,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning","Modular reconfigurable robotic systems enable scalable cooperative surface operations for future lunar missions, yet cooperative cargo transport is hindered by morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. A phase-decomposed reinforcement learning framework is proposed for distributed cooperative units. The task is split into lifting, transportation, and placement, each trained with a dedicated joint-state policy capturing inter-agent coupling. Centralized training stabilizes convergence, while deployment uses onboard proprioception with OptiTrack ground-truth for evaluation. A Markov-state phase controller and failure-sensitive synchronization enforce coordinated, safety-aware halting in real experiments.","Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning  \nAshutosh Mishra∗1, Elian Neppel 1 , Shreya Santra 1 , Antoine Jonquières2 , Muhammad Athallah Naufal3 , Kentaro Uno 1 , and Kazuya Yoshida 1  \narXiv :2607 .00160v1 [ cs .RO] 30 Jun 2026  \nAbstract—Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative cargo transportation remains challenging due to morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. This paper proposesa phase-decomposed reinforcement learning framework for cooperative cargo transport with distributed robotic units. The task is decomposed into lifting, transportation, and placement, each optimized with a dedicated joint-state policy capturing inter-agent coupling. Centralized training promotes stable convergence, while deployment uses onboard proprioception for control and OptiTrack motion capture for ground-truth evaluation and post-processed metrics. A deterministic phase controller expressed in Markov state representation regulates transitions between stages, and a failure-sensitive synchronization mechanism ensures coordinated progression and safetyaware halting during real-world execution. The framework is evaluated in simulation and through controlled field experiments at a JAXA space exploration test facility. Results demonstrate reliable cooperative transport across all stages in both simulation and hardware experiments.  \nI. INTRODUCTION  \nSustained lunar surface missions require robotic systems capable of transporting construction materials, instruments, and logistical payloads under communication delay, terrain uncertainty, and strict safety constraints [1] . Modular reconfigurable robots are particularly suited to such environments, as they enable task-specific morphologies assembled from reusable units, supporting repairability, redundancy, and functional adaptability [2], [3] .  \nModularity introduces configuration-dependent kinematics, actuation topology, and inter-module coupling, reducing policy transferability across morphologies. In cooperative transport, additional coupling arises through the shared payload, producing strong dynamic interactions.  \nExisting decentralized reinforcement learning approaches allow distributed modules to act based on local observations while contributing to global objectives [4], [5] . Centralized training with decentralized execution improves convergence stability in multi-agent settings [6], [7] . Deep reinforcement learning has shown effectiveness for cooperative navigation  \nThis work was supported by JST Moonshot R&D Program, Grant Number JPMJMS223B.  \n1A. Mishra, E. Neppel, S. Santra, K. Uno, and K. Yoshida are with the Space Robotics Lab. (SRL), Department of Aerospace Engineering, Graduate School of Engineering, Tohoku University, Sendai 980–8579, Japan. 2A. Jonquières is with École Centrale de Lille, France. 3M. A. Naufalis with Institut Teknologi Bandung, Indonesia. ∗Corresponding author: A. Mishra ([ashutosh.mishra@dc.tohoku.ac.jp](ashutosh.mishra@dc.tohoku.ac.jp)).  \nFig. 1. System overview of the proposed modular multi-phase reinforcement learning framework. Two distributed wheel arm integrated units cooperatively transport a shared payload through lifting, transportation, and placement stages.  \nand transport [8], [9] . However, most formulations assume fixed topology or optimize a single controller for structurally static teams, without addressing physical reconfiguration and topology-dependent coupling during task execution. For modular systems undergoing physical reconfiguration and multi-stage contact transitions, unified policy optimization can induce gradient interference across heterogeneous dynamics and fails to explicitly model configuration-dependent coupling [10], [2] .  \nIn cooperative lunar cargo transport, the problem is comp","cbCaipERz6QDdnwh","https://ap.wps.com/l/cbCaipERz6QDdnwh","pdf",4932333,1,"English","en",105,"# Introduction\n## Problem Setting and Challenges\n## Proposed Phase-Decomposed Framework\n## Contributions\n## Validation and Evaluation","[{\"question\":\"Why is cooperative lunar cargo transport difficult for distributed modular robots?\",\"answer\":\"It is difficult due to morphology-dependent topology changes, payload-induced dynamic coupling, long-horizon decisions, and strict safety constraints that must be respected during execution.\"},{\"question\":\"How does the proposed method structure the learning problem?\",\"answer\":\"It decomposes the overall task into three phases—lifting, transportation, and placement—optimizing each phase with a dedicated joint-state policy that models inter-agent coupling.\"},{\"question\":\"How does the framework ensure safe and coordinated progression during real execution?\",\"answer\":\"A deterministic phase controller based on a Markov state representation regulates transitions, while a failure-sensitive synchronization mechanism coordinates modules and triggers safety-aware halting when needed.\"}]",1784196469,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"distributed-multi-robot-lunar-cargo-transportation-via-phase-decomposed-reinforcement-learning","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/distributed-multi-robot-lunar-cargo-transportation-via-phase-decomposed-reinforcement-learning/84535/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is cooperative lunar cargo transport difficult for distributed modular robots?","Question",{"text":74,"@type":75},"It is difficult due to morphology-dependent topology changes, payload-induced dynamic coupling, long-horizon decisions, and strict safety constraints that must be respected during execution.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed method structure the learning problem?",{"text":79,"@type":75},"It decomposes the overall task into three phases—lifting, transportation, and placement—optimizing each phase with a dedicated joint-state policy that models inter-agent coupling.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the framework ensure safe and coordinated progression during real execution?",{"text":83,"@type":75},"A deterministic phase controller based on a Markov state representation regulates transitions, while a failure-sensitive synchronization mechanism coordinates modules and triggers safety-aware halting when needed.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]