[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83031-en":3,"doc-seo-83031-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83031,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","RoboTALES Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures","Pretrained video generative models can serve as visuomotor control backbones, yet their imagined futures may drift from task intent and are not reliably action-conditional, limiting planning and policy extraction. RoboTALES introduces a single-stage framework to learn task-aligned simulated futures and use them to train robot policies. It combines an LLM-based hierarchical planner for subgoal-guided imagination with a VLM-based critic that scores rollouts and provides reward feedback. Evaluations on RoboCasa and LIBERO10 show consistent superiority, especially on long-horizon manipulation.","arXiv :2607 .060 18v 1 [ cs .RO] 7 Jul 2026  \nRoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures  \nHanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu, Nicklas Hansen, and  \nManmohan Chandraker  \nUniversity of California, San Diego, CA 92093, USA Correspondence to: [hgani@ucsd.edu](hgani@ucsd.edu)  \nAbstract. Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model’s imagination; and (2) a VLM-based critic that evaluates these “imagined” futures and uses reward-based feedback to keep the model’s internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALESon diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at [https://github.com/hananshafi/RoboTALES](https://github.com/hananshafi/RoboTALES).  \nKeywords: Video Generation · Diffusion Models · Hierarchical Reasoning · Policy Optimization  \n1 Introduction  \nHumans rarely act purely reactively; instead, they decompose goals into ordered subtasks, mentally simulate possible futures, and commit to an action only after rejecting undesirable outcomes [9, 38] . This capacity for hierarchical, goal-directed imagination remains an ongoing challenge for autonomous agents. Agents operating on high-dimensional visual observations must reason over extended horizons, anticipate the consequences of their actions, and select among many plausible futures. While model-free reinforcement learning can achieve strong task performance [20, 39, 48], it offers no explicit reasoning substrate and remains sample-inefficient. Model-based approaches address this by learning a world model that enables an agent to “imagine before acting” [19, 21, 22, 24, 47, 51] .  \nThe idea of leveraging diffusion-based [8, 27, 50] video generation models [11, 13] directly for robot control has recently gained considerable momentum.  \n2 Hanan Gani et al.  \nFig. 1: Overview. An LLM Planner decomposes the task instruction (e.g.,“make acoffee”) into ordered subtasks that condition the video generator to generate future frames for each milestone. A frozen VLM Critic scores these imagined rollouts against the goal instruction, steering the video generator model toward semantically aligned futures. The resulting goal-conditioned features drive the policy to produce precise robot actions for long-horizon manipulation.  \nWorks such as Video-Policy [35], Gen2Act [6], and ViPRA [46] demonstrate that video generators can serve as powerful priors for policy learning, using generated rollouts to either supervise or directly condition action generation. However, video generation alone does not yet constitute a world model for control. This is because the underlying video generators are still optimized primarily for visual realism or reconstruction objectives, hence their simulated futures can remain weakly grounded in task intent; moreover, making generated videos faithfully follow robot actions is itself nontrivial [34] .  \nIn parallel, LLMs have demonstrated a remarkable ability to decompose longhorizon tasks into grounded subtasks and guide robot behavior through natural language [3, 10, 29, 49] . Yet, most prior works treat language and prediction as loosely coupled modules: language influences which action to e","cbCaifuNQ8lZSHcv","https://ap.wps.com/l/cbCaifuNQ8lZSHcv","pdf",12598986,2,1,30,"English","en",105,"# Introduction\n# Overview of RoboTALES\n## LLM Planner and temporal subgoals\n## VLM Critic for semantic faithfulness\n## Evaluation and results","[{\"question\":\"What problem does RoboTALES address in existing video-generative control methods?\",\"answer\":\"It addresses the tendency of imagined futures to drift from task intent and the lack of reliable action-conditional control, which makes planning and policy extraction difficult.\"},{\"question\":\"How does RoboTALES keep simulated futures aligned with the task?\",\"answer\":\"RoboTALES uses an LLM planner to decompose the instruction into ordered subgoals that condition video generation, and a VLM critic that scores the imagined rollouts and provides reward-based feedback to preserve semantic goal faithfulness.\"},{\"question\":\"What evidence is provided that RoboTALES works better than prior approaches?\",\"answer\":\"Experiments on diverse manipulation tasks from RoboCasa and LIBERO10 show that RoboTALES consistently outperforms existing methods, with the biggest gains on long-horizon tasks.\"}]",1784184773,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robotales-learning-reasoning-guided-robot-policies-via-task-aligned-simulated-futures","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/robotales-learning-reasoning-guided-robot-policies-via-task-aligned-simulated-futures/83031/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RoboTALES address in existing video-generative control methods?","Question",{"text":75,"@type":76},"It addresses the tendency of imagined futures to drift from task intent and the lack of reliable action-conditional control, which makes planning and policy extraction difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RoboTALES keep simulated futures aligned with the task?",{"text":80,"@type":76},"RoboTALES uses an LLM planner to decompose the instruction into ordered subgoals that condition video generation, and a VLM critic that scores the imagined rollouts and provides reward-based feedback to preserve semantic goal faithfulness.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is provided that RoboTALES works better than prior approaches?",{"text":84,"@type":76},"Experiments on diverse manipulation tasks from RoboCasa and LIBERO10 show that RoboTALES consistently outperforms existing methods, with the biggest gains on long-horizon tasks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":22,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]