[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84610-en":3,"doc-seo-84610-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84610,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","RoboWorld Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation","Video world models offer a scalable way to evaluate generalist robot policies without physical deployment, but rollout reliability and throughput remain constrained by world-model errors and slow iterative inference. ROBOWORLD introduces an automated evaluation pipeline combining a fast autoregressive video world model with a task-progress-aware vision-language model judge to score rollouts. STEP FORCING improves long-horizon action-conditioned rollouts by reducing train–test context mismatch while preserving action–observation dynamics. Results align strongly with real-world evaluations across tasks and environments, achieving Pearson r=0.989 and Spearman ρ=0.970.","RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation  \nByeongguk Jeon∗ KAIST, Config  \nSeonghyeon Ye∗  \nKAIST  \nJaeHyeok Doo  \nKAIST  \nSungdong Kim  \nKAIST, Config  \nMinjoon Seo  \nKAIST, Config  \nHyungmok Son  \nConfig  \nKimin Lee  \nKAIST, Config  \nProject Website: [https://byeongguks.github.io/RoboWorld/](https://byeongguks.github.io/RoboWorld/)  \narXiv :2607 .01060v2 [ cs .RO] 13 Jul 2026  \nRoboWorld  \nInitial View Interactive Rollout from World Model VLM Scoring  \n. . .  \n. . .  \nGeneralist Robot Policy  \nFigure 1: Overview of ROBOWORLD. ROBOWORLD evaluates robot policies via closed-loop rollouts in a video world model scored by a task-progress-aware VLM judge, yielding rankings that strongly correlate with real-world evaluations.  \nAbstract: Video world models are emerging as a scalable alternative for evaluating generalist robot policies, bypassing the physical constraints and engineering burdens of real-world deployment. However, evaluating policies with video world models remains challenging, as world-model errors can make generated rollouts unreliable and slow inference limits large-scale throughput. We introduce ROBOWORLD, an automated evaluation pipeline that pairs a fast autoregressive video world model with a task-progress-aware vision-language model scoring.  \nTo enable reliable long-horizon autoregressive world-model rollouts, we propose STEP FORCING, which combines anchored and one-step self-forwarded contexts to reduce train–test mismatch while preserving action–observation dynamics. Together, these components enable ROBOWORLD to align strongly with real-world robot evaluation across tasks and environments, achieving Pearson’s r = 0 .989 and Spearman’s ρ = 0 .970.  \nKeywords: Robot Policy Evaluation, World Model  \n1 Introduction  \nReliable and accessible evaluation has significantly accelerated progress in foundation models for vision and language [1, 2, 3, 4] . Generalist robot policies, also referred to as Vision-Language-Action  \n∗Equal contribution.  \nFigure 2: Left Upper: STEP FORCING shares the noise schedule between training and inference.(a) Diffusion Forcing conditions on noisy ground-truth (red) . (b) Self Forcing conditions on selfgenerated context (green) via repeated forward rollouts. (c) STEP FORCING conditions on the onestep self-forwarded prior (green) or anchor step (red) .  \n(VLA) models, have made rapid progress in generalizing across tasks, objects, and environments, requiring numerous rollouts across diverse conditions for reliable evaluation [5, 6, 7, 8, 9, 10, 11, 12] . However, scaling robot policy evaluation in the real world remains challenging. Each rollout requires physical robots and human operators, limiting the number of policies and conditions that can be evaluated. Simulation-based evaluation [13, 14, 15, 16, 17, 18] avoids these costs but requires engineering for asset and environment setup, and sim-to-real gaps compromise its reliability [19, 20] . Recent efforts reduce human intervention [21] or automate real-to-sim environment construction [22, 23], but they still rely on physical robot setups or prior access to target environments.  \nBuilt on pretrained video diffusion models [24, 25, 26, 27, 28, 29, 30, 31], video world models offer a scalable alternative for policy evaluation by enabling interactive closed-loop rollouts [32, 33, 34, 35, 36, 37, 38] . This property makes it possible to evaluate policies in new environments without physical robot setup or extensive simulator engineering. However, two challenges remain:  \n(1) world-model artifacts can corrupt long-horizon rollouts, making evaluation unreliable, and (2) slow inference from iterative denoising processes reduces evaluation throughput. Especially, such artifacts propagate into evaluation outcomes when Vision-Language Model (VLM)-based scoring reduces each rollout to a binary success score [39, 40] .  \nWe introduce ROBOWORLD, a scalable automated evaluation pipeline that combines a","cbCaiuQx7RJ1t6PI","https://ap.wps.com/l/cbCaiuQx7RJ1t6PI","pdf",10075170,2,1,27,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What is STEP FORCING, and why is it important?\",\"answer\":\"STEP FORCING trains the world model to predict clean frames from a few-step self-forwarded prior while using an inference-matched denoising schedule. It reduces train–test context mismatch and maintains action controllability for reliable long-horizon rollouts.\"}]",1784197110,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"roboworld-fast-and-reliable-neural-simulators-for-generalist-robot-policy-evaluation","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/roboworld-fast-and-reliable-neural-simulators-for-generalist-robot-policy-evaluation/84610/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is STEP FORCING, and why is it important?","Question",{"text":75,"@type":76},"STEP FORCING trains the world model to predict clean frames from a few-step self-forwarded prior while using an inference-matched denoising schedule. It reduces train–test context mismatch and maintains action controllability for reliable long-horizon rollouts.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]