[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82159-en":3,"doc-seo-82159-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82159,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Learning More from Less: Reinforcement Learning from Hindsight","Reinforcement learning used for post-training vision-language-action (VLA) robot models suffers from severe sample-efficiency limits because each update requires costly, slow robot rollouts. Manipulation training often uses sparse rewards, so an early weak policy fails almost every commanded instruction and yields little learning signal, even when the robot executes coherent, task-relevant behavior. Learning from Hindsight (LfH) relabels failed rollouts by scoring what the robot actually achieved in language, improving sample efficiency up to 5× on out-of-distribution LIBERO-PRO tasks and transferring to a physical Franka robot.","arXiv :2607 .09042v 1 [ cs .LG] 10 Jul 2026  \nLearning More from Less: Reinforcement Learning from Hindsight  \nIris Xu 1,2,* , Sunshine Jiang 1 , John Marangola 1 , Nitish Dashora 1 , Richard Li 1 , Thomas Liu 1 , Zexue He3 , Yuheng Zhi4 , Alex Pentland3 , Pulkit Agrawal 1 , Zhang-Wei Hong 1,2  \n1Massachusetts Institute of Technology 2MIT-IBM Computing Research Lab 3 Stanford University 4University of California, San Diego  \nAbstract: Reinforcement learning (RL) is increasingly used to post-train visionlanguage-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, so a weak policy fails almost every rollout early in training and has little to learn from, even when those failures execute coherent behavior. Such a failure, however, is a success at a different task.  \nWe present Learning from Hindsight (LfH), which brings hindsight relabeling to RL post-training of VLAs by scoring failed rollouts against the tasks they actually achieved. A single vision-language model relabels both the instruction and the reward, proposing a hindsight instruction for a group of failed rollouts and scoring how well each satisfies it, and the policy trains on the relabeled and original rollouts jointly. Because VLAs generalize across language, relabeling in language lets the policy learn more from the same trajectories. On out-of-distribution LIBERO-PRO tasks, where standard RL improves only slowly, LfH achieves 5 × improvement in sample efficiency, and outperforms a dense progress-reward baseline. The gains hold across VLA backbones and on a physical Franka robot.  \n1 Introduction  \nReinforcement learning (RL) [1] has become a standard tool for post-training large language models [2, 3], and is increasingly being used to fine-tune vision-language-action (VLA) models for robot control [4] . In robotics, however, post-training is constrained by a much harsher data bottleneck: each RL update depends on rollouts collected from a physical robot, which are slow, expensive, and often difficult to scale. Improving sample efficiency is therefore important for making RL post-training practical for VLAs.  \nThis challenge is especially acute in manipulation, where VLAs are commonly fine-tuned with sparse rewards. A rollout receives reward only if the robot completes the commanded instruction, leaving all other behavior uncredited. Sparse rewards are a long-standing obstacle in RL [5, 6]: before a weak policy can improve, it must first discover a successful trajectory through exploration. Early in training, nearly all rollouts therefore appear useless. For example, when instructed to “close the microwave,”the robot may instead pick up a cup (Figure 1) . This behavior is coherent and task-relevant in a broad manipulation sense, but it does not satisfy the commanded instruction. Standard RL treats the entire rollout as a failure, ignoring the fact that the robot successfully executed another meaningful skill.  \nOur key observation is that many failed rollouts are failures only relative to the original instruction. Hindsight relabeling [7, 8, 9, 10, 11] exploits this by evaluating a trajectory against the task it actually achieved rather than the task it was commanded to solve. The cup-picking rollout, for instance, can be relabeled as an instance of “pick up the cup.” This is distinct from dense progress rewards [12], which provide finer feedback for the same commanded instruction. Hindsight changes the task assignment itself, converting an otherwise unrewarded failure into a successful demonstration of a different, related instruction. Since VLAs are conditioned on language and can generalize across instructions,  \n*Correspondence: [irisxu@mit.edu](irisxu@mit.edu)  \n“put the container in the bowl”  \n“move the container”  \nVLA Training:  \nNo learning signal  \nStandard RL  \nSignal for meaningful skills  \nLearning from ","cbCaiceA66752Uv7","https://ap.wps.com/l/cbCaiceA66752Uv7","pdf",2394040,1,23,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does Learning from Hindsight (LfH) target in VLA reinforcement learning post-training?\",\"answer\":\"LfH targets the low sample efficiency caused by slow, costly robot rollouts, especially when manipulation tasks use sparse rewards that provide almost no learning signal early in training.\"},{\"question\":\"How does LfH turn failed rollouts into useful training data?\",\"answer\":\"LfH uses a vision-language model to infer the achieved behavior from rollout observations, then proposes a hindsight instruction and scores how well each failed rollout satisfies it, creating relabeled reward and instruction supervision.\"},{\"question\":\"What performance improvements does LfH report?\",\"answer\":\"On out-of-distribution LIBERO-PRO tasks, LfH improves sample efficiency by 5× versus standard RL, reaches the final success rate in about one fifth of training steps, and also outperforms a dense progress-reward baseline; gains transfer to a physical Franka robot.\"}]",1784178508,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"learning-more-from-less-reinforcement-learning-from-hindsight","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/learning-more-from-less-reinforcement-learning-from-hindsight/82159/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Learning from Hindsight (LfH) target in VLA reinforcement learning post-training?","Question",{"text":75,"@type":76},"LfH targets the low sample efficiency caused by slow, costly robot rollouts, especially when manipulation tasks use sparse rewards that provide almost no learning signal early in training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LfH turn failed rollouts into useful training data?",{"text":80,"@type":76},"LfH uses a vision-language model to infer the achieved behavior from rollout observations, then proposes a hindsight instruction and scores how well each failed rollout satisfies it, creating relabeled reward and instruction supervision.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance improvements does LfH report?",{"text":84,"@type":76},"On out-of-distribution LIBERO-PRO tasks, LfH improves sample efficiency by 5× versus standard RL, reaches the final success rate in about one fifth of training steps, and also outperforms a dense progress-reward baseline; gains transfer to a physical Franka robot.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]