[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82065-en":3,"doc-seo-82065-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82065,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Prompt-Driven Exploration","Exploration is central to reinforcement learning when a policy cannot improve by sampling actions it already prefers. Standard action-noise methods only produce local jitter and rollouts near the original behavior, failing to realize the global perturbations often required to escape weak policies. Prompt-conditioned LLM/VLA models enable global behavior changes through natural-language prompts. Prompt-Driven Exploration refines prompts using a VLM that analyzes rollout videos, diagnoses failures, and rewrites prompts to improve success and sample efficiency.","arXiv :2607 .08837v 1 [ cs .LG] 9 Jul 2026  \nPrompt-Driven Exploration  \nSunshine Jiang 1,3 , John Marangola 1,3 , David Zhang 1,3 , Raghuram Kowdeed3 , Ruiyang Luo 1,3 , Nitish Dashora 1,3 , Richard Li 1,3 , Pulkit Agrawal 1,3 , Zhang-Wei Hong 1,2,3  \nMassachusetts Institute of Technology1 , MIT-IBM Computing Research Lab2 , Improbable AI Lab3  \nAbstract  \nExploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce.  \nLarge language models (LLMs) and vision-language-action (VLA) models offera pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt toelicit better behavior next time. This procedure realizes posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintainsan implicit distribution over useful prompts and updates it from observed rollouts.  \nWe call this strategy Prompt-Driven Exploration (PDE) . Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at [https://xinyunsunshine.github.io/prompt-rl](https://xinyunsunshine.github.io/prompt-rl).  \n1 Introduction  \nReinforcement learning (RL) [20] has become a dominant post-training paradigm for foundation models because it enables scalable self-improvement beyond supervised learning. RL fine-tuning has unlocked strong reasoning and mathematical capabilities in large language models (LLMs) [18, 14] and enabled direct alignment with human preferences in large diffusion-based text-to-image models [6] . However, self-improvement is bottlenecked by exploration: a policy can only surpass its current behavior by producing rollouts different from those it already favors, so that RL can reinforce the higher-reward ones.  \nStandard practice perturbs the policy in action space by sampling actions stochastically rather than greedily [46, 36, 15] . But such action-level noise only jitters individual actions and yields rollouts close to the original; it induces only local exploration [32, 13], and the set of rollouts reachable by step-wise noise shrinks rapidly with the horizon and action dimension. Escaping a weak policy often requires global perturbations that alter behavior across the entire rollout, which action noise cannot produce. This limitation is especially pronounced in settings without a strong warm start, particularly vision-language-action (VLA) model fine-tuning on manipulation, where state-of-the-art models often start at near-zero success rates [22, 5] .  \nIf action noise only produces local exploration, how can we perturb the policy globally? Foundation models offer a pathway. LLMs [7, 1] and VLAs [22, 5] are conditioned on a natural language prompt, and since the entire rollout follows from it, modifying the prompt induces global changes. Figure 1 illustrates this behavior: given the prompt put the green container on the bottom rack,” the policy picks up the container but fails to place it fully on the rack. Action noise merely jitters the arm without changing the strategy. Rephrasing the prompt as put the green container completely on  \nPreprint.  \nOriginal prompt  \n“put the green container on the bottom rack”  \nOptimized prompt  \n“put the green container completely on the bottom rack”  \nFigure 1: Left: Given the prompt “put","cbCaibNBRQerJMnm","https://ap.wps.com/l/cbCaibNBRQerJMnm","pdf",4311732,1,29,"English","en",105,"# Abstract\n# Introduction\n## Exploration in reinforcement learning\n## Limitations of action noise\n## Global exploration via prompts","[{\"question\":\"Why does standard RL exploration with action noise struggle with weak policies?\",\"answer\":\"Action noise only jitters individual actions and tends to yield rollouts close to the original behavior, providing mainly local exploration. Escaping a weak policy often needs global perturbations that action noise cannot reliably produce.\"},{\"question\":\"How does Prompt-Driven Exploration (PDE) achieve global exploration?\",\"answer\":\"PDE conditions the policy on a natural-language prompt, and because the entire rollout follows from the prompt, changing the prompt induces global behavior changes. It then updates prompts based on the outcomes of rollouts.\"},{\"question\":\"How does PDE refine prompts when rewards are sparse?\",\"answer\":\"When success is rare, reward signals are too sparse for direct selection. PDE instead uses a VLM to reason over rollout videos, diagnose how the policy responded, and rewrite the prompt to elicit better behavior next time.\"}]",1784177975,73,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"prompt-driven-exploration","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/prompt-driven-exploration/82065/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does standard RL exploration with action noise struggle with weak policies?","Question",{"text":75,"@type":76},"Action noise only jitters individual actions and tends to yield rollouts close to the original behavior, providing mainly local exploration. Escaping a weak policy often needs global perturbations that action noise cannot reliably produce.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Prompt-Driven Exploration (PDE) achieve global exploration?",{"text":80,"@type":76},"PDE conditions the policy on a natural-language prompt, and because the entire rollout follows from the prompt, changing the prompt induces global behavior changes. It then updates prompts based on the outcomes of rollouts.",{"name":82,"@type":73,"acceptedAnswer":83},"How does PDE refine prompts when rewards are sparse?",{"text":84,"@type":76},"When success is rare, reward signals are too sparse for direct selection. PDE instead uses a VLM to reason over rollout videos, diagnose how the policy responded, and rewrite the prompt to elicit better behavior next time.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]