[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81682-en":3,"doc-seo-81682-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81682,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning","Reinforcement learning typically fixes the temporal structure of reward composition through training, leaving the evolution of motivational priorities underexplored. This work introduces an evolutionary framework that discovers developmental reward schedules by combining three biologically inspired components—agency, novelty, and reactivity—with time-varying weights optimized over training. Experiments on DoorKey-6x6 and KeyCorridorS3R1 compare CMA-ES, xNES, DE, and L-SHADE against an extrinsic baseline and hand-designed methods, yielding task-dependent generalization and insights into how evolution shapes schedule structure.","Evolutionary Discovery of Developmental Reward Schedules in Deep  \nReinforcement Learning  \nAlan Nadelsticher Ruvalcaba 1 ,2  \narXiv :2606 .20858v2 [ cs .LG] 10 Jul 2026  \nAbstract—The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational priorities largely unexplored. In this work, we propose an evolutionary framework for discovering developmental reward schedules, in which three distinct biologically inspired motivational components — agency, novelty, and reactivity —are combined through time-varying weights that dynamically shift over the course of training. Evaluated on two sparsereward MiniGrid tasks: DoorKey-6x6 and KeyCorridorS3R1, our framework compares the generalizability of four evolutionary algorithms: CMA-ES, xNES, DE, and L-SHADE against an extrinsically motivated baseline (our main comparison point), and three additional hand-designed methods. On DoorKey-6x6, all evolved methods outperform the non-evolved baselines, with L-SHADE achieving the best performance — an approximate relative mean improvement of 11.4% over the extrinsic only baseline. On KeyCorridorS3R1, CMA-ES achieves the best overall performance, with the remaining evolved methods showing weaker and less reliable generalization capability compared to the extrinsic only baseline. Interestingly, the discovered schedules diverge from our defined developmental ordering, with novelty consistently emerging as the dominant early signal during training, across both tasks. Collectively, our results position evolutionary optimization as a promising approach for developmental reward schedule discovery in deep reinforcement learning, and suggest that what evolution finds to be optimal in computational settings may differ from what it finds to be optimal in biology. The code for this project can be found at: [https://github.com/alannadels/Evolutionary_](https://github.com/alannadels/Evolutionary_)[RL.git](RL.git.)[.](RL.git.)  \nI. INTRODUCTION  \nIn humans, and other mammals alike [1], the dopaminergic reward system is driven by a shifting ensemble of nonarbitrary, extrinsic and intrinsic motivational signals throughout development: neonates first discover behavior-event contingency [2], infants then develop novelty preferences for stimuli of intermediate complexity [3], adolescents exhibit a hyper-responsive striatal reward system that peaks before prefrontal control matures [4], [5], [6] and, finally, mature goal-directed behavior emerges only once corticostriatal connectivity develops in early adulthood [7] . This progression reflects a functionally optimal strategy for motivational development—one that evolution itself discovered over time.  \n© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.  \n1Alan Nadelsticher Ruvalcaba is with the College of Computing at the Georgia Institute of Technology, Georgia, USA; and 2with the School of Engineering and Applied Science at the University of Pennsylvania, Pennsylvania, USA. Email: [alan.nadelsticher@gatech.edu](alan.nadelsticher@gatech.edu)  \nYet, most reinforcement learning (RL) systems treat the temporal structure of reward composition as a fixed design choice, leaving its biologically and evolutionarily grounded dynamics underexplored.  \nIn this work, we propose an evolutionary approach to developmental reward schedule discovery in RL that directly addresses this gap. Specifically, we seek to discover the functionally optimal developmental reward schedule for an RL agent, given a certain task, and we let evolution discover it—as it did with biological organisms.  \nIn ou","cbCaiex5fFlY08l3","https://ap.wps.com/l/cbCaiex5fFlY08l3","pdf",1580338,4,1,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does the paper address in reinforcement learning reward design?\",\"answer\":\"It addresses the fact that most RL systems keep the temporal structure of reward composition fixed, instead of learning how motivational priorities should develop over training.\"},{\"question\":\"How does the proposed method model developmental reward schedules?\",\"answer\":\"It defines a composite reward function from three biologically inspired components—agency, novelty, and reactivity—and uses time-varying weights optimized by an evolutionary outer loop while training agents with PPO.\"},{\"question\":\"What were the key experimental findings on the two MiniGrid tasks?\",\"answer\":\"On DoorKey-6x6, evolved methods outperform non-evolved baselines, with L-SHADE best (about 11.4% relative mean improvement over the extrinsic-only baseline). On KeyCorridorS3R1, CMA-ES performs best overall, while other evolved methods generalize less reliably than the extrinsic-only baseline.\"}]",1784175388,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"evolutionary-discovery-of-developmental-reward-schedules-in-deep-reinforcement-learning","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/evolutionary-discovery-of-developmental-reward-schedules-in-deep-reinforcement-learning/81682/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address in reinforcement learning reward design?","Question",{"text":74,"@type":75},"It addresses the fact that most RL systems keep the temporal structure of reward composition fixed, instead of learning how motivational priorities should develop over training.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed method model developmental reward schedules?",{"text":79,"@type":75},"It defines a composite reward function from three biologically inspired components—agency, novelty, and reactivity—and uses time-varying weights optimized by an evolutionary outer loop while training agents with PPO.",{"name":81,"@type":72,"acceptedAnswer":82},"What were the key experimental findings on the two MiniGrid tasks?",{"text":83,"@type":75},"On DoorKey-6x6, evolved methods outperform non-evolved baselines, with L-SHADE best (about 11.4% relative mean improvement over the extrinsic-only baseline). On KeyCorridorS3R1, CMA-ES performs best overall, while other evolved methods generalize less reliably than the extrinsic-only baseline.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]