[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84850-en":3,"doc-seo-84850-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84850,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Deep Reinforcement Learning for Dynamic Battery Management of Autonomous Order Pickers","Battery charging for Autonomous Mobile Robots (AMRs) in warehouses is a high-impact operational constraint, directly affecting order processing time and system throughput. This work tackles the dynamic AMR charging problem under stochastic order arrivals, where robots must learn optimal charging decisions instead of relying on fixed heuristics. A PPO-based deep reinforcement learning framework is proposed for multi-block warehouses with fixed charging stations, learning station choice and charging duration while accounting for anticipated queuing. Extensive experiments show up to 6% higher order-completion rates, reduced recharging time, and robustness across warehouse layouts and arrival rates.","arXiv :2607 .05683v 1 [ cs .LG] 6 Jul 2026  \nDeep Reinforcement Learning for Dynamic Battery Management of  \nAutonomous Order Pickers  \nTaniya Shaji 1 , Abhay Sobhanan∗1, and Christof Defryn2  \n1 Indian Institute of Management Bangalore, Bannerghatta Road, Bengaluru 560076, Karnataka, India  \n2 University of Antwerp, Prinsstraat 13, Antwerp 2000, Belgium  \nAbstract  \nBattery charging of Autonomous Mobile Robots (AMRs) in warehouses is a critical operational challenge that heavily impacts both order processing times and throughput. In this study, we address the dynamic AMR charging problem under stochastic order arrivals, where robots must learn optimal charging decisions. Traditional fixed-rule heuristics often prove suboptimal in dynamic environments and fail to account for multi-AMR coordination, leading to severe resource inefficiencies. To overcome these limitations, we propose a Proximal Policy Optimization (PPO)-based Deep Reinforcement Learning (DRL) framework designed for multi-block warehouses with fixed charging stations. Our model dynamically learns two key decisions: charging station selection and optimal charging duration, explicitly accounting for anticipated queuing times at the stations. Extensive numerical experiments benchmark the proposed model against state-of-the-art DRL and traditional heuristic approaches. Results demonstrate that our PPO framework increases order-completion rates by up to 6% compared to the strongest baseline, while significantly reducing the total time dedicated to recharging operations. Furthermore, we validate the model’s robustness across diverse warehouse configurations and stochastic arrival rates. Finally, we interpret the learned DRL policy, offering valuable operational insights into its superiority over standard benchmarks.  \nKeywords: Warehouse logistics, Deep reinforcement learning, Battery Charging, Autonomous Mobile Robots, Machine learning  \n1 Introduction  \nModern fulfillment centers increasingly rely on Autonomous Mobile Robots (AMRs) to automate intra-logistics operations such as the storage, retrieval, and transportation of inventory. To maintain high throughput in these structured warehouse environments, centralized planning software assigns tasks to these robots, which must be executed efficiently under strict time, energy, and capacity  \n∗ Corresponding author. Email: [abhay.sobhanan@iimb.ac.in](abhay.sobhanan@iimb.ac.in)  \nconstraints. The scale of these operations is expanding rapidly; for instance, in 2025, Amazon deployed its one-millionth robotic unit (Dresser, 2025) . Their models, such as Titan and Proteus, navigate the warehouse floor, retrieving orders from specific storage locations and transporting them to designated drop-off points, all strictly coordinated by a centralized control system (Greenawalt, 2025) .  \nThe continuous operation of these AMRs is fundamentally constrained by their limited battery capacities. To maintain power, robots must periodically interrupt their assigned tasks and travel to fixed and often inductive or wireless charging stations (Wiferion) . Commercially used lithium-ironphosphate batteries require significant downtime. For instance, they average 2.7 hours for a full charge (McNulty et al. , 2022), while modern models like KUKA’s KMP 1500P take approximately one hour to replenish from 20% to 80%(KUKA AG) . Consequently, charging protocols critically impact warehouse efficiency. In dense warehouse environments with multiple AMRs and limited charging infrastructure, this creates a complex resource allocation problem. Without coordinated scheduling, multiple robots may seek to charge simultaneously, leading to severe bottlenecks, queuing delays, and an unbalanced utilization of charging stations that ultimately degrades overall system throughput.  \nThe complexity of this charging problem is further amplified in dynamic environments where orders arrive stochastically. In a typical zonal routing setup where individual AMRs are ass","cbCaiqyybVkzGP6c","https://ap.wps.com/l/cbCaiqyybVkzGP6c","pdf",5264468,2,1,35,"English","en",105,"# Introduction\n## Motivation and problem setting\n## Challenges with fixed heuristics and stochastic arrivals\n## Proposed DRL approach and contributions","[{\"question\":\"Which decisions does the PPO-based framework learn?\",\"answer\":\"It learns charging station selection and charging duration, and also supports a dynamic return-to-depot drop-off decision to reduce unnecessary travel.\"}]",1784198800,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"deep-reinforcement-learning-for-dynamic-battery-management-of-autonomous-order-pickers","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/deep-reinforcement-learning-for-dynamic-battery-management-of-autonomous-order-pickers/84850/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Which decisions does the PPO-based framework learn?","Question",{"text":75,"@type":76},"It learns charging station selection and charging duration, and also supports a dynamic return-to-depot drop-off decision to reduce unnecessary travel.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]