[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83201-en":3,"doc-seo-83201-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83201,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators","Robotics increasingly relies on World Models (WMs) to simulate actions in an imagined world and produce safety or success verdicts, but those verdicts are only trustworthy if the WM itself is certified. Video-generation metrics like Fréchet Video Distance (FVD) assess visual realism while overlooking whether the world responds correctly to policy actions. A framework is proposed that adapts VV&A, SOTIF, and scenario-based testing into a levels-of-admissibility ladder (L0–L4), embodiment-agnostic and instantiated for autonomous driving.","Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators  \nChristian Oefinger 1 , Finn Rasmus Schfer 1 , Korbinian Moller 1 , Mattia Piccinini 1 , and Johannes Betz 1  \n1Autonomous Vehicle Systems Lab  \nTechnical University of Munich, Garching b. M¨unchen, Germany  \n[Email: christian.oefinger@tum.de](Email: christian.oefinger@tum.de)  \narXiv :2607 .07 196v 1 [ cs .RO] 8 Jul 2026  \nAbstract—Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world, and returning a successor safety verdict. Yet a verdict is only as trustworthy as the WM that produced it, and the WM itself needs to be certified. In video-generation WMs, fidelity metrics such as Frchet Video Distance (FVD) reward visual realism, but ignore whether the world responds correctly to the policy’s actions, including those unseen in training. Classical simulation-based validation assumesa trusted simulator evaluating an untrusted policy, whereas generative WMs are themselves unverified learned artifacts. Hence, we argue that any WM used as a test oracle must first be accredited before its verdicts can serve as evidence. Building on credibility practices from safety-critical simulation, including Verification, Validation & Accreditation (VV&A), Safety of the Intended Functionality (SOTIF), and scenario-based testing standards, we define an admissibility ladder (L0–L4) that a WM must climb before its closed-loop verdicts are accepted as assurance evidence. Our framework is embodiment-agnostic, and is instantiated in autonomous driving (AD), where assurance methods for traditional simulation are most mature. Applied to two driving WMs, the lower rungs reveal a reversal: the model that ranks higher on visual generation quality (L0) ranks lower on action-following (L1–L2), so visual fidelity does not predict the action-robustness a closed-loop verdict depends on.  \nI. INTRODUCTION  \nWMs are generative models that learn an internal representation of an environment’s dynamics, letting them predict how the world responds to an agent’s actions [15, 16] . Recent video-generation models turn this prediction into high-fidelity, controllable simulation [12, 18], with the potential to reshape how robotic systems are developed and tested. Across robotics, the role of WMs is shifting from imagining plausible futures to serving as closed-loop simulators that test the action policies that act within them [17]: A WM rolls out the consequences of a policy in an imagined world, and returns a verdict on success or safety. Thus, WMs are used as test oracles in several robotic embodiments. WMs serve for instance as (i) a testbed for manipulation policies [31], (ii) a benchmark for world-model planning on legged robots [39], or (iii) a generative simulator for closed-loop evaluation in autonomous driving [43] . However, the reliability of WMs is typically left unverified, treating their verdicts as evidence of real-world behavior, without any basis to assess admissibility.  \nTrusting the verdict assumes a high-fidelity imagined world can substitute for reality. For video-generation WMs, which this paper targets rather than reconstruction-based simulators,  \nFig. 1. The levels-of-admissibility ladder (L0–L4). A generative WM used as a closed-loop test oracle earns the right to have its verdict counted as assurance evidence only by climbing the ladder. Verdict validity first appears at L2 . L0–L1 support no admissibility claim, and all guarantees hold only within the declared operating envelope. Table I lists the evidence required at each level.  \nvisual quality alone does not determine task success [9, 45] . Distribution-level video scores, such as the FVD [38], reward perceptual and temporal realism while ignoring physical plausibility and the world’s realistic reactions to the policy’s actions. This disconnect has been demonstrated across embodiments. For legged robots, th","cbCaijz3ZgEz6GPb","https://ap.wps.com/l/cbCaijz3ZgEz6GPb","pdf",300931,1,10,"English","en",105,"# Introduction\n## World Models as closed-loop test oracles\n## Limits of existing fidelity metrics\n## Need to validate the WM as well as the policy\n## Proposed admissibility standard (L0–L4)","[{\"question\":\"Why can a World Model’s verdict be untrustworthy?\",\"answer\":\"A WM’s verdict assumes the imagined world can stand in for reality. For generative video WMs, visual quality metrics do not guarantee that the world responds correctly to the tested policy’s actions.\"},{\"question\":\"How does the proposed approach ensure WM verdicts are evidence?\",\"answer\":\"It requires the WM to climb an admissibility ladder (L0–L4) before its closed-loop verdicts are accepted as assurance evidence. The criteria are derived from VV\\u0026A, SOTIF, and scenario-based testing practices for safety-critical simulation.\"},{\"question\":\"What does the framework reveal when applied to autonomous driving?\",\"answer\":\"When instantiated in autonomous driving on two driving WMs, higher visual generation quality at lower levels (e.g., L0) can coincide with weaker action-following performance at higher admissibility requirements (L1–L2), showing visual fidelity does not predict action robustness.\"}]",1784185924,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"validate-the-dream-before-you-trust-its-verdict-admissibility-for-world-model-simulators","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/validate-the-dream-before-you-trust-its-verdict-admissibility-for-world-model-simulators/83201/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can a World Model’s verdict be untrustworthy?","Question",{"text":75,"@type":76},"A WM’s verdict assumes the imagined world can stand in for reality. For generative video WMs, visual quality metrics do not guarantee that the world responds correctly to the tested policy’s actions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed approach ensure WM verdicts are evidence?",{"text":80,"@type":76},"It requires the WM to climb an admissibility ladder (L0–L4) before its closed-loop verdicts are accepted as assurance evidence. The criteria are derived from VV&A, SOTIF, and scenario-based testing practices for safety-critical simulation.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the framework reveal when applied to autonomous driving?",{"text":84,"@type":76},"When instantiated in autonomous driving on two driving WMs, higher visual generation quality at lower levels (e.g., L0) can coincide with weaker action-following performance at higher admissibility requirements (L1–L2), showing visual fidelity does not predict action robustness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]